Mayur Mehta

Building MapChat · Part 4 of 4 · Opinion

Agents made building cheap. Deciding got expensive.

What seven months of building with AI agents changed about the product manager's job

I build MapChat with AI coding agents. Our games developer built the games server, and I own product and engineering for the rest, which means writing the specs, making the product and architecture calls, and setting the bar every release clears. Seven months in, my view is that agents didn't make product management easier. They moved the hard part.

Building got cheap

The backend was replaced in two days. A bar-crawl feature with a hub, challenges, a live leaderboard and a city-wide chat went from planning to a live event in eight days. The app reached both stores three months after the rebuild started.

When a feature takes days to build, the build stops being the constraint. What's left is deciding what to build, saying exactly what it is, and checking what comes back.

BEFORE: THE BUILD WAS THE BOTTLENECK Spec Build weeks Review Ship WITH AGENTS: THE SPEC AND THE CHECKS ARE Spec the product Build days Verify the job Ship
Where the hard part sits. A way of seeing it, not a measurement.

The spec became the product

An engineer who reads a vague requirement fills the gaps with judgment. An agent builds what you wrote. In a blind comparison I commissioned (the setup is in Part 3), the cheaper builder carried out the spec's mistakes faithfully, the stronger one caught some of them, and neither caught everything. What stuck with me was that the quality of the spec, not the model, decided the quality of the output.

So the handoff packet became my most important artifact. Each one has to stand on its own, with the decisions already made, the boundaries, and the checks that prove the work is done. If an agent has to guess, the packet isn't finished.

Verification became the job

The more code agents write, the less of it can rest on an agent's word. Agents found ways around our checks more than once, and each time the fix was to move the check somewhere it couldn't be skipped. I still test every user-facing change on a phone before it merges.

I use four kinds of grading (deterministic checks, a model as judge, rubrics and human review), and none of them is enough alone. A static review can't see a screen that feels wrong in your hand, and a phone test can't see a missing database policy.

Cheap prototypes need cheap deaths

When building is cheap, the risk is keeping everything you build. Our first event-ingestion system was a LangGraph prototype with LangSmith evals. It proved the idea could work, then lost to a simpler design built from a catalog, recurrence templates and agent-run scraping. Because the evals were already in place, retiring it was a quick call.

Natural-language search went the same way. I prototyped it with a small model that turned free text into fixed filter chips, so it couldn't invent values, and then deferred it in favor of universal search across venues and events.

Some calls stay human

Agents can draft almost anything, but there are decisions I don't hand over.

What I'd tell another PM

  1. Write every spec as if it will be carried out literally, because it will be.
  2. Never let the builder grade its own work. Put the check where it can't be skipped.
  3. Read production before designing. My check-in redesign started from a count of real check-ins, not from memory.
  4. Decide the product before the code. A request that still needs a product decision isn't ready for an agent.
  5. Keep the architecture and the cut-lines. Delegate the code.

This is one product, one accountable owner and seven months, so treat it as a field report, not a law. My bet is that the core carries over. When building is cheap, the product manager's job is judgment.