When a team asks us to help with AI in go-to-market, they almost always ask for the same thing: better messaging, generated automatically, personalized at scale.
It's the wrong end of the system to start from.
We've built a number of these now — audience definition, sourcing, enrichment, scoring, generation, tracking — and the generation step is consistently the least interesting part of the machine. It's also the part that gets all the attention, which is why so many AI outbound efforts produce personalized emails about nothing.
What actually determines whether it works
Take a concrete version of the problem: find companies worth talking to, figure out which ones are actually in market, and reach out with something worth reading.
The generated message is the last four percent of that. Everything upstream decides whether the message can be good at all:
Who you're looking at. A precisely worded email to a badly chosen list is a precisely worded failure. Most of the lift we've ever measured came from tightening the definition of the audience, not from improving the copy.
What you know about them. "Personalization" is a data problem wearing a copywriting costume. A model can only reference what you've actually collected. If your enrichment is a job title and a company size, you get an email that references a job title and a company size — which is exactly the email everyone now deletes on sight, because it's what everyone else's model also produces.
Whether you can tell in-market from merely plausible. This is where the real edge is, and it's almost entirely non-generative: a hiring pattern, a product change, a page that suddenly exists, a team that just grew, a technology choice that implies the problem you solve. Building the pipes to notice those things reliably is unglamorous engineering work with a very high ceiling.
Whether you record what happened. Systems that don't capture outcomes can't improve. Half the "AI GTM" setups we've looked at are open-loop: they generate, they send, and nothing about the next batch is different because of the last one.
Why teams skip it anyway
Because the boring layer doesn't demo.
You can show a founder an email the model wrote in four seconds and get a reaction. You can't easily demo the enrichment pipeline that makes the email worth sending, even though that's the asset. So the budget goes to the visible part, and six weeks later the reply rate looks like the reply rate always did.
There's a second reason, which is that the boring layer is where the actual difficulty lives. Data is missing, inconsistent, rate-limited, and occasionally wrong in ways that are hard to detect. Generation, by comparison, is a solved-ish problem you can rent. Teams gravitate to the part where progress feels fast.
What we do instead
We work backwards from the decision the system exists to make.
If the decision is should a human spend twenty minutes on this account, then the system's job is to be right about that, and the email is downstream. So we spend the early part of an engagement on:
- A crisp definition of the target. Written down, argued about, narrow enough to be wrong.
- The signals that separate in-market from plausible. Usually three to five, and usually specific to the business in a way no generic tool can replicate. This is the part that compounds, because it's built on your data and your observed outcomes, not on a model everyone else can also call.
- Scoring you can inspect. Not a black-box number. A rep needs to see why an account surfaced, or they won't work it — and their disagreement is the most valuable feedback in the system.
- A closed loop. What happened to every record. Without this the system can't learn, and neither can you.
Only then generation — which, when the layers underneath it are right, becomes easy. There's nothing clever left to do, because you already know something true and specific about the account. The model is just writing it down.
The general version
This pattern isn't about go-to-market. We keep finding it everywhere.
The advantage almost never comes from the model. Everyone has the same models, and the gap between them keeps narrowing. The advantage comes from having connected a model to proprietary data, real workflows, and a feedback loop nobody else has.
Which means the durable work in most AI deployments is data and plumbing work. That's an uncomfortable thing to sell and an easy thing to skip, and it's the reason so many AI projects produce a demo and then a plateau.
If you're evaluating an AI initiative inside your company, the useful question isn't which model are we using. It's:
What do we know that our competitors don't, is it captured anywhere a system can reach, and does anything we build get told whether it was right?
If the answer to the last part is no, fix that before you build anything else. It's the least exciting thing on the roadmap and it decides everything above it.