AI/ML Engineer for Large-Scale URL Analysis

Job ID: 39628934

Budget: $250 – $750 AUD

We are looking for a highly skilled AI/ML engineer with strong Python experience to help analyze a large dataset (~4.5 million URLs stored in MongoDB) and build a system that can detect structural patterns and sequential logic in URL formats across various domains (e.g., /product/12345 → /product/12346).

This will support predictive takedown automation and broader content monitoring strategies across e-commerce platforms and marketplaces.

Key Responsibilities:
Analyze large-scale URL datasets stored in MongoDB.

Group and cluster URLs by domain.

Design and implement models to extract and generalize structural URL patterns.

Detect and predict sequential patterns (e.g., numeric or token-based increments).

Build or fine-tune models using LLMs (e.g., T5, FLAN-T5, Mistral) or custom LSTM/CNN architectures.

Develop tools to output patterns in regex or templated format (e.g., /product/{id}).

Optimize pipeline for performance and scalability (batch processing, streaming optional).

Document logic and recommend improvements to enrich pattern recognition over time.

Ideal Skillset:
Strong Python skills (especially in data wrangling and NLP).

Experience with MongoDB (PyMongo, aggregation pipelines).

Familiarity with Hugging Face Transformers (T5, FLAN-T5, Mistral, etc.).

Experience building or fine-tuning LSTM, Transformer, or Autoencoder models.

Regex generation/mining tools (e.g., refex, AutoRegex, regexgen).

Clustering algorithms and similarity measures (Levenshtein, n-gram, FAISS).

Familiarity with web scraping or URL structure analysis is a plus.

Comfortable working with large datasets (millions of rows).

Tools & Frameworks You Might Use:
Python (Pandas, Numpy, Scikit-learn, PyTorch/TensorFlow)

Hugging Face Transformers

MongoDB (via PyMongo)

Regex and pattern mining libraries (re, refex)

Clustering: HDBSCAN, KMeans, Levenshtein

Optional: FAISS, spaCy, NLTK

Deliverables:
Script(s) to analyze and extract patterns from MongoDB-stored URLs.

ML model(s) or heuristic tools that can generalize and detect URL structures.

Regex or templated pattern outputs per domain.

Documentation of logic, decisions, and next steps.

(Optional): Prototype tool to visualize URL pattern clusters or predictions.

Please include in your application:
Links to similar work (GitHub, Kaggle, or private projects)

Short explanation of how you’d approach detecting URL structure patterns

Your availability and estimated timeframe to complete phase 1 (MVP)