The Product Agent

From zero discovery to customer evidence in every brief: requirements validated 60% faster

Head of Product Design · CourtReserve · 2025–2026 Racquet club and facility management SaaS.

The situation

I joined CourtReserve as its first ever design hire, one of a handful of leaders brought in right after the company’s first round of funding to help it scale and put distance between encroaching competitors. Product development ran through a black-box engineering org: the CTO was the single point of contact for the rest of the business on any work requested, and design meant UI work outsourced to an agency. There were already eight product managers. I was told the company had that many because it was trying to build faster and needed more help defining projects to keep engineering busy.

The new VP of Product and I looked at the staff we had and restructured product development into cross-functional squads, each with a PM, a dev lead, and a UX lead.

Eight squads. Eight PMs, most of them early in the discipline. No research practice. No design system. And budget for three designers.

I made a version of the headcount case. But, even a successful one lands six to twelve months out and says nothing about the interim. The usual alternative is triage: put designers on the squads shipping the most visible work and let the rest run uncovered. That splits quality into two tiers which get harder to reconcile every quarter.

So I tried another approach: could we get customer evidence and design decisions to reach a squad as infrastructure rather than as a person’s time? A PM holding good evidence and good components is better than a PM waiting in a queue for design.

This is what that turned into.

Product agent flow-selection

What we built

The Product Agent is a set of interoperating agentic skills that a PM, UX lead, or ops person invokes conversationally. It reads and writes across multiple connected tools. Ours leveraged Linear, Slack, Supabase, Google Drive, Intercom, UserEcho, and Gamma. It covers the lifecycle from the first question through release feedback: research capture and synthesis, discovery planning, brief writing, work breakdown, demo compilation, and launch planning. Sixteen skills shipped. Every PM used it, across 30+ projects.

It was co-architected with our VP of Product. He outlined the planning and delivery half. I outlined the research and evidence half and I built the agentic skills to power it all.

Underneath it, I built a research practice from scratch. Over the course of my time there, the squads generated 100+ customer interviews, 12 advisory board sessions, two surveys at 200+ respondents each, and a continuous read on thousands of support conversations. I also put designers in front of customers at two user conferences. Alongside it, I built a design system: 50+ desktop and mobile React/TypeScript components built on Ant Design, documented in Storybook, adopted by all eight squads.

research-pipeline-3d

The Design Problem

Research had never been unpopular at CourtReserve. It had been skippable, and under a weekly cadence anything skippable gets skipped.

So the target was not making research accessible. Accessible loses to instinct every time, because instinct is free. The target was removing the choice. A PM does not request evidence and then wait on it. They ask for a brief, and the brief arrives with the evidence already pulled, themed, and cited. Research stopped being a step someone decides to take and became a property of the artifact they already wanted. Going on instinct was still possible, but it now meant actively stripping the evidence out rather than simply not getting around to it.

That reframes the work. Most of the interesting decisions were not about what the system could do. They were about what it was forbidden to do, because a system that quietly guesses is worse than no system at all. A PM who trusts a fabricated answer makes a worse decision than a PM who knows they are guessing.

Where the system defers

The first diagram shows where people decide. This is what it took to keep those slots from filling themselves in.

Human judgment gets a protected slot, and the model may not fill it. Themes carry an organic confidence score derived from how much evidence supports them. They also carry a human signal score, set by a person who was in the room and thought something mattered more than the volume suggested. Ranking is the composite, so a medium-confidence theme with strong human signal outranks a high-confidence theme with none. The rule that makes it work: the model must never generate, suggest, propose, or infer a human signal flag. That input exists specifically to encode what the transcript did not capture. A model filling it in would be recording its own inference as a person’s judgment, which corrupts exactly the signal the field was built to hold.

The council probes and never approves. Before an idea becomes a project, five perspectives push on it and surface blind spots. None of them endorses, greenlights, or validates the direction. A PM leaves with sharper questions, not permission to build. The decision stays where it belongs.

Citations resolve to records, not links. Every evidence claim in a brief has to resolve to an actual row in the research or feedback store, with a theme identifier and a source URL. A shared document link is reference material for a human, never just a citation. Attribution without a link is not attribution. It is what makes an evidence-backed brief checkable rather than merely confident. Anyone can follow a claim back to the interview that produced it.

Citing a theme is not the same as addressing it. We track which customer themes have a campaign working on them. A brief citing a theme as evidence does not create that link. Only a person explicitly asserting it does. The distinction sounds pedantic until you watch a coverage map fill itself in with work nobody is actually doing.

We cut the feature that would have automated it. The first version of that coverage map had an automatic matcher pairing themes to campaigns by keyword and goal alignment. At any threshold defensible enough to trust, it fired on roughly one theme in fifteen. That was a ceiling in the shape of the data, not a tuning problem. Shipped, it would have shown leadership a coverage map that looked complete and was not. We cut it and made coverage human-written.

Conflict gets flagged, not resolved. The discovery planner scores an initiative across seven factors and recommends a research depth. When a PM’s brief budgets less time than the score calls for, the system does not silently rescope and does not overrule. It plans to the PM’s number and names the gap. The PM owns the call. The system owns telling them what they are trading.

Capture happens at the source, and there is a gate on it. Before any conversation is logged, the system asks for the recording and will not proceed to a manual write-up until a person confirms there isn’t one. A research layer built on write-ups holds one person’s interpretation. A layer built on transcripts holds what the customer said. Only one of those survives being queried a year later by someone who was not in the room.

A brief cannot ship to no one. No unit of work is final without a named release audience: internal, a named early access group, or specific named customers. General availability is not a valid answer. This is the smallest rule in the system and it closes the loop, because a named audience is what makes it possible to go back and ask how it landed.

evidence-gate-v6 (1)

What that cost

The choice was right for the constraint. It was not free.

Craft depth in some surfaces. The design system did work a designer would otherwise have done. That produces consistent and correct. It does not produce distinctive. In a few places distinctiveness would have been worth more than consistency, and we did not get it.

Partnership. Some squads never had sustained design partnership. They had evidence and they had components. That is not the same thing as having someone in the room, and I would not pretend otherwise to anyone who worked in those squads.

Standing maintenance. A research layer nobody curates degrades into a search index over stale interviews. Choosing leverage means owning it permanently, and that has to be staffed like any other line of work.

What changed

Every product brief carried customer evidence, because the workflow would not advance without it. That is a different thing from teams valuing research. They valued it before. They skipped it anyway.

Time to a validated requirement dropped from weeks to days.

Adoption is the number I would actually point to. Every squad on the design system. Every PM on the Product Agent, across 30+ projects.

What I'd do differently

I identified the measurement gap early and said it out loud. I still got the sequence wrong.

My view then and now: validation tends to happen either before you build or after you ship, and “after” only works if you’ve built the instrumentation to measure it. Most teams that choose “after” haven’t. CourtReserve hadn’t. There was no measurement foundation at all, which meant “we’ll learn it in production” wasn’t a strategy so much as a hope.

I was right about that, and I argued for the validation layer first. In hindsight I should have built the instrumentation layer first. Same systems, same conviction, different order, and a completely different argument. Building post-ship measurement first would have made me the person accelerating the speed agenda rather than the person advocating for the more deliberate model. It’s much easier to make the case for pre-development validation when you can show, with numbers, what shipping without it costs.

That’s the lesson I carry into the next org: lead with the instrumentation. Earn the argument with evidence rather than winning it on principle.