AI machine learning Expert Needed

Job ID: 39534734

Budget: $1,500 – $3,000 USD

Naisaa.AI
We are developing Naisaa.AI, a comprehensive AI content detection platform that identifies whether content (text, images, audio, or video) has been generated or manipulated by AI. The application must provide accurate detection, model attribution, explainability, trust scoring, and user-friendly interfaces via web and API.
We are seeking a quotation for the end-to-end development of the platform, including frontend, backend, machine learning models, deployment, and documentation.
Core Objectives
• Accept and analyzetext, image, audio, and video content.
• Detect AI-generated content using a combination of machine learning models, statistical analysis, and provenance tools.
• Return detailed results: detection outcome, confidence score, explanation, and possible source model (e.g., GPT-4, Midjourney, ElevenLabs).
• Provide a SaaS-style web interface, user account system, and optional file scan history.
• Provide a REST API for 3rd-party integrations (LMS, HR systems, publishers).
• Ensure a scalable, modular, and secure architecture, ideally containerized via Docker.

Functional Requirements: Multi-Modal Detection Modules
1. Text Detection
• Detect AI-generated text from models like GPT-4, Claude, etc.
• Use perplexity, burstiness, and fine-tuned transformer model (e.g., RoBERTa).
• Return result with explanation and heatmap.
• Estimate which model family likely generated the text.
2. Image Detection
• Detect AI-generated images from tools like Midjourney, DALL•E, Stable Diffusion.
• Use CNN or ViT model trained on GAN features.
• Analyze image metadata (EXIF, C2PA watermarking).
• Provide visual explanation overlay if possible.
3. Audio Detection
• Detect voice cloning or synthetic speech (e.g., ElevenLabs, TTS models).
• Analyze MFCCs, jitter, pitch, waveform consistency.
• Compare with human baselines or embeddings (e.g., Resemblyzer).
• Return voice confidence score + waveform/spectrogram diagnostics.
4. Video Detection
• Detect deepfakes and synthetic videos.
• Frame-level facial mapping and inconsistency analysis.
• Analyze compression artifacts, facial morphing, lip-sync discrepancies.
• Use pretrained models (e.g., XceptionNet, DeepFaceLab) or integrate open-source detectors.

Platform Features
Web Application (Frontend)
• Built in React + Tailwind CSS
• Upload or paste content (text, image, audio, video)
• View results with clear trust score and explanation
• Login, dashboard, scan history
Backend + API
• FastAPI or Node.js backend with modular endpoints:
/detect-text, /detect-image, /detect-audio, /detect-video
• Auth (JWT or Firebase)
• PostgreSQL for user & scan data
• S3 or GCP Bucket for file storage
• Redis or task queue for async processing (e.g., video/audio jobs)
Admin + Monitoring
• Admin interface (scan logs, user management, usage stats)
• Logging and basic analytics
• Rate limiting / abuse prevention
AI/ML Development
• Implement or fine-tune:
o Transformer for text classification
o CNN/ViT for image classification
o RNN/CNN for audio detection
o Deepfake classifiers for video

• Train on custom and open datasets
• Integrate SHAP or LIME for explainability
• Compute and return trust scores and model attribution probabilities
DevOps / Infrastructure
• Full Docker support (docker-compose)
• Environment support (.env)
• Deployment-ready (Render, Railway, AWS, GCP — your recommendation)
• Secure storage and API architecture
• Optional: CI/CD pipeline and unit tests
Deliverables
• Full source code (frontend, backend, ML modules)
• Deployment scripts or Docker containers
• README and module-by-module documentation
• Trained model checkpoints
• Admin credentials + test accounts
• Setup of a staging server (optional)
Quote Request
Please include:
1. Total cost estimate (Fixed price) ?
2. Timeframe, estimated total timeline and milestone breakdown?
3. Team & Technology Stack: Languages, ML frameworks, databases, hosting…?
Add-ons to quote separately:
• Mobile responsiveness / mobile app
• LMS/HR system plugins
• Video watermark verification (C2PA full support)

TEAM:
1. AI/ML Lead
Goal: Build and fine-tune detection models for text, image, audio, and video.
Skills:
• Expert in NLP (transformers, GPT) and vision/audio models
• Hands-on with PyTorch, HuggingFace, TensorFlow, Librosa
• Experience with deepfake detection and voice cloning analysis
• Can build explainability layers (SHAP, LIME)
• Trains models with real/synthetic data + watermark detection (e.g., C2PA)
2. Full-Stack Developer
Goal: Build and integrate the frontend, backend, APIs, and user dashboards.
Skills:
• FastAPI / Node.js backend experience
• React + Tailwind for modern UI
• Authentication (JWT, Firebase)
• Experience integrating file uploads, audio/video scanning workflows
• RESTful API design and testing (Postman, Swagger)

3. DevOps Engineer
Goal: Build and maintain a secure, scalable, and automated deployment pipeline.
Skills:
• Docker + docker-compose
• CI/CD with GitHub Actions, GitLab, or CircleCI
• Cloud deployments (Render, Railway, AWS, or GCP)
• GPU instance configuration for model inference
• API key and secret management, monitoring (Sentry, Prometheus)

4. UI/UX Designer
Goal: Design a smooth, intuitive interface for scanning, visualizing results, and understanding AI risk.
Skills:
• Figma, Adobe XD, or Sketch
• UX for dashboards, trust scoring, and explainability overlays
• Designs for desktop + Chrome Extension
• Familiar with SaaS UI patterns (auth, onboarding, scan history, API dashboard)

API Integration Engineer
Goal: Create plugins/connectors for LMS, HR platforms, and Chrome.
Skills:
• Familiar with LTI 1.3, LMS APIs (Moodle, Canvas)
• Chrome Extension development (JS/React)
• Webhooks + 3rd-party API integration (Zapier, Workday, Greenhouse)
• Plugin packaging and documentation