Most of what we ship under this heading is retrieval, not model work. A client has knowledge spread across PDFs, a wiki, a ticketing system and three databases, and wants answers that are correct and traceable to a source. The interesting engineering is in chunking, retrieval quality, permission filtering and evaluation — not in prompt wording.
You would own that end to end: ingestion through to the evaluation set that tells us whether a change made answers better or just different.
What you would do
- Design ingestion and chunking for real client corpora — inconsistent PDFs, tables, scanned documents, half-maintained wikis.
- Build retrieval that filters by identity in the same query as the search, so a user never retrieves a document they cannot see.
- Create evaluation sets before shipping, so answer quality is a measurement rather than an impression.
- Keep inference cost predictable — caching, per-account caps, and picking the smallest model that clears the bar.
- Design against prompt injection from the start: scoped tools, confirmation on irreversible actions, provenance on every chunk.
What we are looking for
- Production Python or TypeScript, and comfort in a relational database beyond an ORM.
- You have shipped something using an LLM API to real users, and can describe what went wrong with it.
- Working knowledge of embeddings and vector search — pgvector or equivalent.
- You measure before optimising, and can say what 'better' meant on a system you improved.
- Clear written English. Most decisions here are argued in writing.