The problem
[FILL IN: What was broken or slow, who felt it and what it cost. For example: a four-person support team answered the same shipment questions all day, and response times slipped every peak season.]
[What had been tried before, and why it didn’t work.]
The approach
[How we scoped it: which workflows we looked at, what we chose not to automate, and how success was measured.]
- 01Shadowed the team and labelled [N] real tickets
- 02Built an eval set before writing a single prompt
- 03Shipped a draft-only version to two agents in week [N]
What we built
[The system in plain language: what comes in, what the AI does, where a human stays in the loop, and what goes out.]
[Diagram, screenshot or short video of the system]
Results
[What changed, with the numbers from the outcomes row and how they were measured. Include one human detail: what the team does with the time now.]
Stack used
- Claude
- Python
- Postgres + pgvector
- [Helpdesk tool]
- Vercel
- [Eval tooling]