Designing an AI agent companies trust with their customer data
Ops teams spend their days fixing customer records and answering requests. I lead design on LeanData's AI agent that does that work for them, and changes nothing without a person's approval.
Bad data quietly costs millions in revenue
Leads route to the wrong rep. Deals go cold with no owner. Account health slips and nobody notices.
The cleanup never ends, and nobody owns it.
The queue is illustrative. The numbers come from partner calls and our own request history.
cases/month in one interviewed team's RevOps queue
of our own requests were a problem we had seen before - fixed once and came back again
of open pipeline sat with owners who had already left, on our own org
before anyone noticed an enrichment contract had lapsed and data had stopped flowing
how one partner now sees bulk fixes, after a single update set off automations that broke their CRM
Not a data cleaner. An agent that works the whole queue, ahead of you
Plenty of tools detect bad data. We bet on one that also answers requests and does the work, which made one question the center of the product: what is it allowed to do alone?
Every request from Slack, Jira and Salesforce Cases ends in an answer, a plan, or a clear route-out.
Problems surface before a rep, a forecast review or an auditor finds them.
The routine half runs behind approval gates, with an audit trail and a way back.
Policies and past decisions live in the product, not in one person's head.
A teammate said this, and a partner said the same thing. Teams can build their own agent now. What's hard to build is the part that makes it safe to use: their policies, approvals, and a record of every change. That's the part I'm designing, and it can't feel heavy.
A Monday morning with the agent
Each moment is a design decision about when the agent works quietly and when it stops and asks.
You arrive to decisions, not a backlog
The one decision that needs you, and why the agent stopped.
Records fixed, new issues and plans in flight, at a glance.
Bad data in dollars, the number leaders ask about first.
It hears a request in Slack, then stops at your rules
A rep flags a likely acquisition. The agent finds both accounts.
Entity match, parent and child structure, and the policies it pulled.
Above the approval bar, the merge waits for a person.
We first called these tickets. We're moving to requests (the final name is still open) so no one thinks we're replacing their ticketing system. We sync with tools like Jira, and requests feed in from there.
It finds what nobody reported
What the agent can't decide ranks above what it is handling.
Sorted by what a problem costs, not how many records it touches.
Each finding shows its trend since the last sweep.

A finding is a condition. It closes only when a re-scan confirms the fix
I mapped every state before we argued about any of them. The board became what we ran reviews against.
I wanted the agent's solo work in a separate log. A teammate pointed out that means checking two places. Now it's one list, with Handled hidden by default.
It asks up to four questions, not ten
412 titles match no role, so it asks what should happen to them.
Each option says what happens. The agent's pick is marked.
Something else is always an option.
Early versions asked 8 to 10 questions before doing anything. People pasted the data into Claude instead to get a faster answer. Now the agent asks only what it can't infer, four at most, each with a suggested answer.
You see what moves before anything runs
Routing graphs, buying groups and journeys this run will touch.
Low-confidence titles wait for review instead of guessing.
Checks after every batch. Every lead can roll back for 30 days.
Partners don't want to approve every run. They want to pick which kinds of changes wait for them and which the agent can just make.
Every change can explain itself
Open any record the agent touched to see what changed, why, and who approved it.
It handles work no workflow was built for
A rep gave notice. Help me plan her transition.
Nobody wrote down how books get split, so it learns from past handoffs.
Step one keeps new work from reaching her today. Renewals wait for your call.
The product needed a system before it needed screens
With no time to start from scratch, I used shadcn and Tailwind as the base, so tokens map cleanly to code. Then I wrote the layer that made it ours.

Build, show it to customers, change it, repeat
We didn't research first and design second. We started from our own mess, put something in front of people early, and kept listening while we built.
internal RevOps requests, read one by one. 29% just stopped with no outcome.
distinct use cases from 15 companies. Organizational changes and governance came up most.
design partners we met every two weeks. We gave the most weight to needs several of them raised on their own.
Between sessions, feedback surfaced at the next call or not at all. So I built a feedback button that posts straight to our Slack with the page it came from. When I didn't have an answer, I wrote the open question down as a position and built something the team could react to.

More design-ready work than capacity. So I shipped the frontend
Rather than let UI queue behind feature work, I pushed PRs directly against the product. Findings, the Data Dictionary, knowledge docs, the feedback button and the nav all shipped from my branches.


We are heading into early access and I'm still designing, building and iterating as feedback keeps coming in. I'll keep this case study updated as it grows, but take a peek below at my learnings so far + what's next!
AI is changing how we work, not just what we build
Building an AI product with AI tools is changing how I work. So far, two lessons have stuck out.
If you made it, you own it
With AI, I did more reading on this project than ever. Docs, specs and artifacts multiplied, and most of them came as long prose. Generating is cheap now. Reading isn't.
So if my name is on an artifact, I know exactly what's in it, I stand behind it, and it's on me to make it readable and short.
So far, the longest artifact I received for this project so far was 34,030 words, which would take an average reader 2.5 hours to read.
Roles are bleeding into each other
AI lowered the cost of stepping into someone else's lane. I pushed frontend PRs and did PM work, strategizing and scoping. My engineers and PM took on prototyping too.
General availability in December
This is a running case study. The question we measure against stays the same: would an admin have done what the agent proposed?













