AI-Driven Procurement Classification Optimization

Job ID: 39711954

Budget: $10,000 – $20,000 USD

Machine Learning Engineer - NICE Classification System Migration

Project Overview:
We’re seeking an experienced Machine Learning Engineer to migrate our existing LLM-based government procurement classification system to a proprietary, high-performance ML solution. This project involves working with 8.1+ million classified records to build an autonomous classification model that will reduce costs by 98% and increase processing speed by 1000x.

Technical Challenge
- Current State: LLM-based system processing procurement data at $0.01/classification with 2-second latency
- Target State: Native ML model achieving <50ms classification time at $0.0002/classification with >95% accuracy

Key Responsibilities

Phase 1: ML Model Development (4 weeks)
- Data Engineering: Clean and preprocess 8M+ records from PostgreSQL database
- Feature Engineering: Implement TF-IDF vectorization with n-gram analysis
- Model Training: Develop Random Forest classifier with hyperparameter optimization
- Validation: Achieve >95% accuracy through cross-validation against existing LLM results
- API Development: Build FastAPI classification service with batch processing capabilities

Phase 2: System Architecture (3 weeks)
- Database Design: Create normalized product structure (Joinsy.Products table)
- ETL Pipeline: Implement automated data extraction and transformation
- Performance Optimization: Ensure 1000+ classifications/minute throughput
- Monitoring: Integrate logging, confidence metrics, and anomaly detection

Phase 3: Advanced Analytics (4 weeks)
- Geoeconomic Analysis: Build regional price comparison algorithms
- Seasonality Detection: Implement time-series analysis for demand patterns
- Predictive Modeling: Develop price forecasting and trend analysis capabilities
- Dashboard Integration: Create executive-level analytics interface

Required Technical Skills

Core ML/AI Expertise:
- Machine Learning: Random Forest, SVM, ensemble methods
- NLP/Text Processing: TF-IDF, n-grams, text classification
- Feature Engineering: Text preprocessing, dimensionality reduction
- Model Validation: Cross-validation, hyperparameter tuning, performance metrics

Technical Stack:
- Python: scikit-learn, pandas, numpy, FastAPI
- Database: PostgreSQL, SQL optimization
- Cloud/DevOps: Docker, API deployment, monitoring
- Analytics: Time-series analysis, statistical modeling

Experience Requirements:
- 5+ years in machine learning engineering
- Proven experience with text classification at scale (1M+ records)
- Production ML systems deployment and optimization
- Government/procurement data experience (preferred)
- Portuguese/English language processing capabilities

Project Specifications

Performance Targets:
- Speed: <50ms per classification (vs. current 2000ms)
- Accuracy: >95% classification accuracy
- Throughput: 1000+ classifications per minute
- Cost: 98% reduction in operational costs

Deliverables:
1. Trained ML Model with documented accuracy metrics
2. REST API Service with comprehensive documentation
3. Database Schema for normalized product structure
4. Analytics Pipeline for geoeconomic and temporal analysis
5. Performance Dashboard with key business metrics
6. Technical Documentation and deployment guides

Data Assets Available:
- 8,116,598 classified records in PostgreSQL
- NICE classification taxonomy with 45 categories
- Existing LLM outputs for model validation
- Government procurement data across multiple regions (including Brazil)

Budget & Timeline:
- Duration: 13 weeks (3.25 months)
- Budget Range: $15,000 - $25,000 USD
- Payment Schedule: Milestone-based payments
- Work Style: Remote with weekly progress meetings

Ideal Candidate Profile - You’re perfect for this project if you:
- Have built production text classification systems handling millions of records
- Enjoy optimizing ML models for performance and cost efficiency
- Can work independently while maintaining clear communication
- Have experience with government/enterprise data processing
- Understand the importance of model interpretability and compliance

Application Requirements

Please include in your proposal:
1. Relevant Portfolio: Similar text classification projects at scale
2. Technical Approach: High-level strategy for model development
3. Timeline Breakdown: Detailed milestone schedule
4. Questions: Any clarifications needed about the data or requirements

Success Metrics: This project will be considered successful when we achieve:
- 98% cost reduction in classification operations
- 50x improvement in processing throughput
- >95% accuracy maintained or improved from current system
- Production-ready deployment with monitoring and alerting

Ready to transform government procurement analysis with cutting-edge ML?

We’re looking for an engineer who can turn 8+ million data points into a lightning-fast, cost-effective classification system that will revolutionize how we analyze public spending patterns.