Computer Vision Engineer for PropTech Company Ideally office based or Hybrid

Job ID: 40566374

Budget: £18 – £36 GBP

Lead Engineer — Computer Vision & LLM Systems
London (hybrid/in-office preferred) | Full-time**

## The company
We are an AI-powered property intelligence platform. We turn a Rightmove or Zoopla listing into a full development appraisal — GDV, build costs, planning probability, extendable envelope — derived directly from floor plan geometry and plot data. Our browser extension and AI assistant, Otto, put this in front of UK estate agents at the point of instruction; a consumer-facing product is in build. We are mid-pilot with Glue Dog, a major UK independent estate agent network. If you need process, roadmap certainty, or a large team around you before you can ship, this isn't the role.

## The role

You will own two things end to end: the computer vision pipeline that turns raw floor plans and plot data into structured geometry, and the LLM/RAG architecture behind Otto. You will also carry production ownership — you are the person who gets paged when the appraisal engine returns a wrong number to a paying agent mid-pilot, and the person who designs the system so that happens as rarely as possible.

You will report directly to the founder and work alongside two engineers (frontend/backend generalists) whom you will technically lead on anything touching CV or LLM infrastructure.

### What you'll actually be doing in month one
- Auditing and hardening the current floor-plan-to-geometry pipeline: room boundary extraction, wall/door/window detection, area computation from scanned and vector inputs.
- Closing a live architectural vulnerability in Otto's consumer tier — full appraisal data is currently passed into LLM context with only prompt-level suppression; this needs to become an architectural (not prompt-level) guarantee before consumer launch.
- Reviewing and rebuilding as needed the retrieval-and-narrate architecture: locked JSON tool calls against a deterministic calculation engine, with the LLM producing zero independent arithmetic.
- Fixing data pipeline fragility — e.g. postcode-passing failures that silently fall back to national averages instead of failing loudly.

### Ongoing
- Extend the CV thesis: deriving extendable development envelope from floor plan geometry plus plot constraints (permitted development rules, setbacks, massing) — this is the core differentiated IP of the product, and largely still needs to be built out, not just maintained.
- Own uptime, monitoring, and incident response for the calculation engine and Otto during the Glue Dog pilot and beyond. You build the observability stack; you also carry the pager until we can hire under you.
- Design and enforce output constraints on the LLM layer consistent with FCA/FSMA compliance requirements (no ROI%/yield/annualised-return framing, RICS development-appraisal terminology only) — the compliance logic is defined; you own making it structurally unbreakable rather than instruction-dependent.
- Mentor and technically direct the two existing engineers on infrastructure and ML-adjacent work.

## Requirements

**Computer vision**
- Production experience with object detection/segmentation (YOLO, Detectron2, Mask R-CNN, or transformer-based segmentation) applied to structured/technical imagery — floor plans, CAD, architectural drawings, or comparable (satellite/aerial, medical, industrial).
- Geometric reasoning on top of CV output: polygon extraction, area computation, constraint-based spatial reasoning.
- Experience with OCR on low-quality scanned/PDF inputs, and building pipelines that degrade gracefully rather than silently producing wrong answers.

**LLM/RAG**
- Production experience building tool-calling LLM architectures with hard output constraints — not prototype chatbots. You should already understand why prompt-level suppression is not a security boundary.
- RAG pipeline experience: embeddings, vector DB (ChromaDB, Pinecone, Weaviate, or equivalent), retrieval tuning.
- Comfort building schema-enforced, denylist-constrained generation for a regulated domain.

**Otto-specific**
- Context isolation between subscription tiers within a single LLM session — the candidate must be able to explain, unprompted, why passing full paid-tier data into a shared context window and relying on prompt instructions to withhold it from free-tier users is an architectural failure, not a prompt-engineering problem, and must have implemented tenant/permission-scoped retrieval before (row-level security, scoped embeddings, or separate context construction per access tier) rather than a single retrieval call filtered after the fact.
- Structured output enforcement against a fixed schema (JSON schema validation, function-calling with strict argument typing, or grammar-constrained decoding) sufficient to guarantee Otto cannot emit a number it did not receive from the deterministic calculation engine — this is a harder constraint than typical RAG citation-grounding and the candidate should have shipped systems where a malformed or out-of-schema model response is rejected and retried rather than passed to the user.
- Adversarial testing / red-teaming of LLM outputs specifically for regulatory leakage — designing test suites that attempt to get the model to produce ROI%, yield, annualised return, or other FSMA s.21-exposed language despite denylist and system-prompt controls, and iterating the architecture (not just the prompt) until those attempts fail.
- Experience building or maintaining a denylist/allowlist enforcement layer that sits outside the model call (post-generation validation, not just pre-generation instruction) so that compliance failures are caught even when the model doesn't follow instructions.
- Familiarity with retrieval-and-narrate patterns specifically — i.e., has built systems where the LLM's only job is to fetch a pre-computed answer and phrase it, with zero tolerance for the model performing its own arithmetic, unit conversion, or inference on numeric fields.

**Infrastructure**
- Real on-call experience: you have carried a pager, triaged a production incident, written a postmortem.
- Observability (Datadog/Grafana/Sentry or equivalent), structured logging, alerting.
- Production-grade Python backend (FastAPI/Flask), Docker, AWS, PostgreSQL.

**Seniority**
- 5+ years, with at least one full cycle of shipping *and operating* an ML system that had real users and real failure modes. Kaggle competitions and demos don't substitute for this.
- Comfortable operating with significant ambiguity and no safety net.

**Nice to have, not required**
- Geospatial/GIS experience (PostGIS, OS MasterMap, QGIS).
- Prior proptech exposure.
- UK planning/RICS domain familiarity.
- FCA/FSMA-adjacent compliance experience.

## Practical
- British National (Overseas) / right-to-work in the UK required; we cannot sponsor at this stage.
- Compensation: competitive salary, cash only — no equity. Open to a fractional/contract arrangement (3–6 months) for the right person if full-time isn't the right fit yet.