SENIOR PYTHON / AI AGENT ENGINEER
Confidential employer
Senior Python AI Agent Developer (LangGraph, RAG, pgvector)
TYPE OF WORK: Full Time
SALARY: $ ---------- USD per month, based on demonstrated experience
HOURS PER WEEK: 40-50
BEFORE YOU APPLY
Put your one-paragraph answer to Question 1 at the very top of your message. Applications that do not start with that answer are deleted unread. We are not being difficult, we just get a lot of copy-paste applications and this is the fastest way to find the people who actually do this work.
ABOUT HEARTSTAMP
HeartStamp is a US AI-native greeting card platform. People create a one-of-a-kind card with us and we print and mail it to the door, or deliver it as an animated digital card. We launched in the US in 2026 and we are shipping fast. We are a small distributed team and every engineer owns a real part of the product.
THE ROLE
We are hiring a Senior Python AI Agent Developer to work on Stampy, our conversational card assistant, with a clear path to owning it end to end.
Stampy is not a simple chatbot. It is a stateful, tool-using agent that understands what a customer wants, searches a live catalogue of tens of thousands of cards, drives image generation, and hands off into a print-ready pipeline. It is the product.
The hard part of this job is not getting a model to answer. It is making the answer reliable. Did the captured intent survive the handoff. Did the right tool get called with the right value. Does the flow recover when the model drifts. Can you prove any of it with an evaluation suite instead of guessing. If you have shipped agents to real users, you already know that the failures that hurt are the confident wrong answers, not the crashes.
You will work directly with our Tech Lead and Principal Engineer, and you will own decisions, not just tickets.
THE HARD PART
Read this section carefully. It is the real job and it is what we will talk about if we speak.
A conversational agent and a product catalogue speak different languages. The agent can express far more than any storefront has pages for, and the translation between those two vocabularies is where agent products quietly break. Not with errors. With plausible answers that are wrong.
The layer that turns what a customer means into catalogue state, and keeps both vocabularies honest with each other, is the highest-value thing you will own here. Getting it right is mostly not a model problem. It is a contract problem, a state problem, and an evaluation problem.
If that sounds interesting rather than tedious, we should talk.
WHAT YOU WILL DO
The agent layer
- Build and maintain stateful, tool-using LLM workflows in LangChain and LangGraph: structured outputs, routing, checkpoints and memory, approval gates, multi-step flows.
- Own the deterministic layer around the model: reliable slot and state capture, stage gating, guarding tool calls, and graceful recovery when the model goes off track.
- Integrate multiple LLM providers (OpenAI, Google Gemini, Vertex AI, OpenRouter, xAI) with model selection, prompt versioning, rate limit handling and fallbacks. We are model agnostic on purpose.
- Manage interactive performance: per-turn latency budgets, response streaming, prompt caching, model routing for sub-tasks, and cost versus speed trade-offs.
Retrieval
- Build and tune our RAG pipeline on pgvector: embeddings, HNSW indexing, chunking and metadata design, hybrid retrieval, ranking, relevance tuning, and fallback strategies.
- Retrieval that serves a real catalogue with real inventory, not a document question and answer demo.
The resolution layer
- The mapping between conversational intent and catalogue state: occasion slugs, recipient and filter tokens, style tokens, and the routes they resolve to.
- Keeping the agent's vocabulary and the storefront's vocabulary in sync, as a checked contract rather than a runtime guess.
- Round-trip guarantees: if we put a value in a URL, reading it back gives you the choice that produced it.
Proving it works
- pytest, async tests, mocked external services, and agent evaluation suites.
- Golden datasets, automated behavioral tests, retrieval metrics, LLM-as-a-judge, and production feedback loops.
- Structured logging, tracing, error monitoring, and LLM tracing with LangSmith.
You will also work in
- FastAPI: async APIs, Pydantic, REST design, auth, error handling, integration tests.
- PostgreSQL with SQLAlchemy or SQLModel, and query performance.
- Celery and Redis for background AI pipelines: retries, idempotency, backfills, long-running jobs. We have people who own the infrastructure, so you do not need to be deep here, but you need to be comfortable.
REQUIREMENTS
- At least 5 years of professional Python engineering, including owning production services.
- At least 2 years building LLM features, agent workflows, semantic search, or RAG systems that real customers used.
- Real experience shipping retrieval on PostgreSQL and pgvector, not only prototype chatbots.
- Solid FastAPI experience inside a distributed system of APIs, workers, queues and third party services.
- Experience evaluating AI output quality using golden datasets, behavioral tests, retrieval metrics, or LLM-as-a-judge.
- Good judgment on reliability: provider outages, model failures, retries, graceful degradation, and feature-flagged rollouts.
- Comfortable with Docker, CI/CD, and AWS.
- Strong written English. You will be arguing about architecture in Slack with senior engineers and writing design notes people act on.
- Able to overlap at least 4 hours daily with 9am to 5pm US Eastern Time. This is required, not preferred.
NICE TO HAVE
- Enough TypeScript to follow a bug across the boundary. Some of our worst defects live half in Python and half in the frontend filter state. You do not need to be a React developer, but you need to be able to read it.
- Multimodal AI, especially image generation or image understanding pipelines. Strongly preferred.
- DeepEval, LangSmith evaluations, Sentry, or PostHog.
- S3-compatible storage, image moderation, or content safety.
- A consumer product where AI output directly drives a purchase or a creative result.
WHAT THIS ROLE IS NOT
- Not a prompt writing role. We already have a prompt engineer and a card library team.
- Not a research role. What you build ships to customers this quarter.
- Not a rebuild from scratch. The platform is live. You are strengthening and extending it.
- Not an infrastructure role. We have a DevOps engineer and a principal engineer who own that.
SCREENING QUESTIONS
Short and specific beats long and general. Applications without answers will not be reviewed.
1. An agent captures the user's intent correctly, the value is present and correct in the request, and the user still gets the wrong result. Nothing errors and all tests pass. How do you find this, and what check do you add so it cannot happen again?
2. Describe an evaluation suite you built for an agent or retrieval system. What did it assert, how often did it run, and what did it catch that a unit test would not?
3. You have a multi-step conversation where the model sometimes skips a required step. How do you make it reliable without making it feel scripted?
4. Describe a retrieval system you shipped on PostgreSQL and pgvector. What was the corpus, how did you chunk it, and how did you know the ranking was good?
5. What are your current working hours, and what overlap can you commit to with 9am to 5pm US Eastern?
HOW TO APPLY
Answer to Question 1 at the very top of your message. Then your resume, a link to production work you can talk about in detail, and your answers to the rest.
We review daily and move fast. Strong candidates get a short paid trial task: a small, self-contained problem drawn from the kind of work you would actually be doing. About two hours, paid at your rate, and it is the closest thing to the actual job we can show you before we commit.
Track this external role
Sign in to save this listing and apply through the original source.
Sign in to SaveReport JobOpen Original Listing