Loading...
In the Field · A series from real jobs

I run customer service. Here's the system I built — one piece at a time.

No engineer, no budget request, no waiting on a roadmap. Just a customer experience manager, a help-desk platform, and a coding agent — describing what I needed in plain English until it existed.

Editor's note This is one person's account — a customer experience manager in e-commerce — shared with us and lightly edited, with their name withheld by request. It's the first in In the Field — real, step-by-step stories from people in different jobs, showing exactly what they built. Their words begin now.
The short version

A customer experience manager turned a pile of scattered conversations into one system — one that tracks the team's efficiency, grades every customer conversation for quality, and reaches out to customers before they ask. All of it built by describing it in plain English. No code written by hand, no software purchased.

Get onto one platform Automate the reports Pull every transcript Grade every conversation Reach out proactively
Every conversation
read and reviewed — not just a small sample
A month → next day
how fast coaching now reaches agents
Under $1 a day
to run the whole quality side — about $300 a year

Managing customer experience throws a lot of numbers at you, and if you can't see them clearly, you're already in trouble. How fast are we answering? Are people waiting on hold? Is one channel quietly falling apart while another looks fine? You're accountable for all of it — and most days you're guessing.

Out of the box, a support platform gives you some pre-built reports with the basics. But it's not tailored to your business, and it looks different on every channel — voice, SMS, live chat, email, social. If you don't have one source of truth for how your team is actually doing, you're at a disadvantage before you start. This is the story of closing that gap, in the order it actually happened — and the honest version includes the part a coding agent didn't do.

One thing worth watching as you read: a coding agent wrote every step below. But that doesn't mean there's AI running inside every step. There are two kinds, and the tags tell you which is which:

Plain code The agent wrote the code — then stepped out. What runs every night is just code, doing the same thing the same way. No AI inside it anymore.
Code + AI judgment Here the AI is part of the daily run — reading conversations and making a call no fixed rule could. That's the judgment layer →
0
First, the groundwork Not Claude Code

Get everything onto one platform

Before any of the clever stuff, I had to fix the foundation: our data lived in too many places. So we moved everything onto Gladly, which gave us a single home for every conversation across every channel. I want to be straight about this — that move had nothing to do with a coding agent. It was a normal platform decision, and it's worth saying out loud, because not every win here came from code.

Gladly

A genuine shout-out to GladlyHonestly, Gladly is a fantastic platform and I can't recommend it enough. Most help desks treat every contact as a fresh ticket; Gladly builds the whole thing around the customer — one lifelong conversation across voice, SMS, chat, email, and social, instead of a pile of disconnected tickets. That single-customer view is the thing that made everything later in this story possible. If your conversations are scattered across tools, fixing that first is the highest-leverage move you can make — and Gladly is a genuinely great place to land.

And if you're reading this from a different setup: you don't have to switch platforms to do any of what follows. The same approach works with whatever software you already pay for — as long as it'll let a program reach its data (most will), an agent can connect to it. Getting everything onto one platform first isn't required; it's just a big head start, because every tool you build after this only has one place to look.

But here's the catch I hit immediately: even with one great platform, the pre-built reports still didn't show me my team the way I needed to see it. The data was finally in one place. Seeing it my way was a different problem — and that's where Claude Code came in.

1
Step one · Visibility Plain code

Let the reports build themselves

I asked Claude Code to build a connection into Gladly, then pointed it at Gladly's documentation and told it to learn its way around. In plain terms: I gave it permission to reach into our own account and pull our data automatically, and handed it the manual for how — so I'd never have to log in and export anything by hand again.

In plain English — "connection" and "documentation" Most serious software has a side door built for other programs (an API): hand it the right key — a credential the platform gives you — and a program can read your data without a human clicking around. And every side door comes with a manual (the documentation) that lists exactly what you can ask for. I didn't read it; I told the agent to. It went through the manual, found every report we had, and worked out how to pull each one. APIs and pulling data, in plain English →

Once it could find the reports, I had it replicate them — download all the Gladly data I cared about (not everything, just what's relevant to how we run) — and then do that automatically, every night. Suddenly I had a complete, current picture of our operation in one place.

From there I built the reports I actually wanted: our service level, each agent's service level by hour, when people were taking breaks, efficiency hour by hour — basically everything you'd want from an operations standpoint. It pulled what had been seven or eight separate Gladly reports into one tailored view, and put it in front of the people making the calls on scheduling and hiring.

The reporting I used to assemble by hand now just exists — current the moment I open it, no one keying anything in, no waiting on someone to spin up an export.

Here's the part worth being clear about: Claude Code wrote this, but there's no AI in it. Once the program existed, the AI stepped out — what runs every night is plain code, doing the exact same thing the same way, with nothing thinking or deciding. For reporting, that's not a limitation, it's the whole point: numbers you can trust precisely because nothing is improvising. Why that distinction matters →

2
Step two · The raw material Plain code

Get every transcript out, not just the numbers

With efficiency handled, the obvious missing piece was quality. Anyone in customer experience knows the grind of quality assurance: it's a full-time job, you're wading through an ocean of conversations, and you know you're only ever seeing a sliver. You want to feel sure your agents are treating customers the way you'd treat them yourself — and spot-checking a handful a day doesn't give you that.

I already had the connection into Gladly from step one, so I asked Claude Code whether we could pull the actual conversation transcripts, too. It went back through the documentation, tested the options, and worked out how to download transcripts for every channel — voice, email, chat, all of it.

Then we had a back-and-forth to throw out what didn't belong: spam, automated marketplace messages, anything that had nothing to do with quality. Once that was clean, it built a job to download all of our real transcripts every night, right alongside the reports.

3
Step three · Quality Code + AI judgment

Have every single conversation read and graded — overnight

This is the step that changed the job. We built a script that goes through those transcripts every day after they download and grades each one against a rubric specific to our business — a rubric I built out with Claude Code's help, in plain English, until it matched what "good" actually means for us.

This is the other kind of step — the one where the AI doesn't just write the code, it is part of the daily run. Every single day it reads every conversation and decides whether a supervisor needs to look at it, based on our criteria. That read isn't keyword-matching — it's a real call on how the conversation went, the kind only judgment can make. This is the "AI" part, and here's what that really means →

It narrows hundreds of conversations a day down to fewer than ten — and now every conversation a supervisor reviews is actually worth their time.

The knock-on effect was bigger than I expected. Every conversation gets seen now, not the tiny sample we used to catch. And because the coaching moments land the very next morning instead of whenever someone got to them, our training loop went from something like a month down to the next day. Agents get better faster, and customers feel it sooner.

4
Step four · Above and beyond Code + AI judgment

Reach out before customers ask — and let it grade its own calls

Once efficiency and quality were both handled, I had something valuable sitting there: every conversation, in one place. So I asked it to go further. A script reads through those transcripts looking for moments worth following up on — someone mentioned they needed the product for a birthday, a wedding, a big event.

And because the connection into Gladly already existed from step one, it doesn't just find those moments — it acts on them. It goes back into Gladly and creates a task on the right day, scheduled for an agent, with the context of why we're reaching out and a suggestion of what to say.

It doesn't stop there. It reads how customers respond to that outreach and grades its own judgment — was that a good moment to reach out? — and gets sharper from real conversations over time.

That last part is the piece I still find a little wild. The same system that does the work also checks its own calls against what actually happened, and improves. I didn't write any of it by hand. I described what I wanted, in the language I'd use to explain it to a new hire.

Where this leaves me

Efficiency, covered. Quality, covered — every conversation, not a sample. And on top of that, outreach that's actually proactive and gets better on its own. None of it was a product we bought; all of it fits our business exactly, because I built it to.

And the honest order matters. Step zero — getting onto one good platform — wasn't a coding-agent thing at all. But everything after it, the part that turned "data in one place" into a system that runs our operation, came from being able to build. I got to decide, piece by piece, which jobs wanted a reliable, same-every-time program (the reporting) and which needed real judgment (the quality read, the outreach) — each tool right for its job, neither one lesser. That choice used to belong to whoever had an engineering team. Now it's mine.

under $1 / day
≈ $300 a year
And the part that still surprises people when I tell them: the AI doing all that judgment — reading and grading every single conversation, every day — costs us less than a dollar a day to run. Call it around $300 a year. The reporting side is a reliable program, so it runs for next to nothing. That's the whole quality operation, for about what one mid-tier software subscription costs in a month.

This is one of a growing series. See more from In the Field → · New to all this? Start from the beginning →
Built something like this in your job? Share your story → it helps the next person.