In this article:
Want us to find IT vendors for you?
Share your vendor requirements with one of our account managers, then we build a vetted shortlist and arrange introductory calls with each vendor.
Book a call

How to Build a Sales Coaching System That Scales

How to build a sales coaching system that scales: why coaching breaks down as teams grow, how to set a standard backed by evidence, and how to measure it.

Author:
Date

Sales coaching is the ongoing development of a seller's ability to run better customer conversations — through specific, repeated feedback tied to what actually happened in those conversations.

The usual explanation for why it disappoints is that managers don't have enough time. That's true, and it isn't the binding constraint. An organization with unlimited manager hours and no agreed standard will coach inconsistently at greater volume.

Coaching doesn't fail on effort. It fails on evidence.

Here is the question almost nobody asks before building a coaching system, and the one this article is organized around: what standard is your coaching measured against, and where did that standard come from?

What is sales coaching?

Sales coaching is the ongoing development of a seller's ability to run better customer conversations, delivered through specific and repeated feedback anchored in real conversations with buyers rather than general advice.

Three things get used interchangeably here, and separating them is worth doing properly because the confusion is expensive.

Training transfers knowledge. It happens at onboarding, at a kickoff, at a product launch. It's one-directional and it's finite. Training is how a seller learns what your product does and which methodology you run.

Pipeline review inspects deals. It asks what's going to close, what's slipping, and what the forecast is. It's necessary, and it is not coaching, though it consumes most of the time nominally allocated to coaching.

Coaching develops capability. It asks what this seller needs to do differently in the next conversation, and it works on one thing at a time until that thing changes.

Most of the published guidance in this category makes the first two distinctions well. Almost none of it goes to the next question, which is the one that determines whether any of it works: coaching develops capability toward what? Better than last month? Better than the rep at the next desk? Better than some absolute definition of a good conversation — and if so, whose definition, derived from what?

Hold that question. Everything below depends on it.

Why sales coaching doesn't scale

Coaching works in small teams and degrades as organizations grow. That's not a coincidence or a management failure. It's three structural problems that only appear at scale.

It runs on memory. A manager coaches from the conversations they happened to sit in — a small, non-random sample skewed toward deals that were already in trouble. Whatever the manager didn't see doesn't get coached, and what they did see is remembered imperfectly by the time the one-to-one happens. This is where "be more consultative" comes from. It's what's left when specifics have faded.

It relies on the wrong standard. Most organizations use intelligent platforms or tools these days to coach from data. These tools are nice but they use your own company's data to create a standard. So if your organization has a mediocre sales system, the standard becomes mediocre too.

It's manager-dependent. Two sellers on the same team, working the same accounts, receive different standards depending on who they report to. One manager cares about discovery. Another cares about next-step discipline. A third coaches whatever they were good at when they carried a bag. None of them are wrong, exactly, and that's the problem — when a standard lives in a person, it leaves when the person does. This is what "inconsistent coaching" actually means, and it's why coaching programs evaporate after a reorganization.

It's untethered. Because nobody wrote down what good looks like in advance, nobody can say afterward whether the coaching worked. The organization has spent real hours and can produce no evidence of what changed. That absence is why coaching budgets are the first thing questioned in a difficult quarter.

Notice that none of the three is solved by more time. Give that organization twice the manager capacity and you get twice the volume of memory-based, manager-dependent, untethered coaching. Time is a real constraint. It's the second one.

What standard is your coaching measured against?

Here is where I think the category has a genuine blind spot, and I say that as someone building a product in it.

Open any sales coaching guide and you'll find a scorecard. 25-point rubrics. 7-category templates. 5-star scales across discovery, objection handling, and next steps. They are thoughtfully constructed and they are almost always asserted — presented as self-evidently correct, with no account of where the thresholds came from.

Ask the question directly of whatever scorecard your organization uses. Why is the talk-time threshold that number and not five points either side of it? Why does discovery quality sit at that bar? What evidence connects those specific thresholds to conversations that became revenue?

For most teams the honest answer is one of three: it came from a vendor's default template, it came from someone senior's experience, or it came from analyzing the organization's own best calls.

The third one sounds the most rigorous. It's the one I want to spend time on, because it contains a trap that is very hard to see from inside.

The ceiling problem

A platform that analyzes your company's conversations can tell you a great deal. It can show you what your top performers do differently from your bottom performers, surface the moments worth reviewing, and give you coverage no manager could achieve manually. That is real and valuable, and I don't want to diminish it.

But notice what that platform is comparing against. It benchmarks your sellers against other sellers in your company. The standard it produces is your own distribution — and specifically, its upper end.

Which means: if your organization's discovery is weak across the board, a benchmark drawn from your best conversations certifies that weakness as excellence. Your top rep becomes the target. The team converges on them. And the system reports improvement the entire time, because relative to your internal distribution, improvement is exactly what happened.

You cannot detect this from inside. A team coaching itself toward its own ceiling and a team coaching itself toward a genuine standard produce identical dashboards. Both show variance narrowing. Both show reps approaching the benchmark. Only one of them is getting good.

This is not a criticism of conversation intelligence platforms, and I want to be careful here because it would be easy to read it as one. It's a statement about vantage point.

Gong, Chorus, and every platform built on a customer's own conversations are structurally confined to that customer's data — that is what makes them valuable, and it is also what makes a cross-category standard unavailable to them. No amount of engineering changes it. A tool that sees one company's conversations, however deeply, cannot tell that company what good looks like across the market it sells into.

Establishing that requires having watched a lot of different companies sell to the same buyers.

Score your own scorecard

Not your reps. Your standard. The clearer your answers are, the clearer will be the assessment on where you stand.

1 Which criteria does your team coach against?

Select the ones your managers actually use — not the ones you’d like them to.

Up to three of your own.

If you'd rather not use the tool, the audit takes ten minutes on paper. List every criterion your managers coach against. Next to each, write the threshold as a number. Next to that, write where the number came from. Most teams find one of two patterns: criteria with no numbers attached, which means each manager is applying their own, or numbers with no source, which means nobody can defend them when a seller pushes back. Both are fixable. Neither is visible until you write it down.

How AI changed sales coaching, and what it didn't fix

Something genuinely improved in the last few years, and it's worth being precise about what.

Coverage improved. Manual review reaches a sliver of a team's conversations — whatever a manager can listen to between everything else. Automated analysis reaches all of them. That is a real change and I wouldn't want to go back. Sampling bias in coaching was a serious problem and it is largely solved.

The standard didn't improve. Scoring is now cheap. What scoring is compared against is exactly as well-evidenced as it was before, which for most organizations means a rubric someone wrote based on judgment and experience.

Put those together and you get the thing to watch for: an organization can now score one hundred percent of its conversations against an unvalidated standard.

That's just measurement of the wrong thing, delivered faster and with more confidence than before, because a number that came out of a system feels more objective than a number that came out of a manager's opinion. It usually isn't. It's the same opinion, encoded once and applied consistently.

Consistency is genuinely worth something — a consistently applied standard is better than five inconsistent ones, even if the standard is imperfect. But it is not the same as a standard that's right, and organizations routinely mistake the first for the second.

The constraint was never analysis capacity. It was that nobody could say what the analysis should be compared to.

How to measure sales coaching effectiveness

If you take one operational thing from this article, take this section. Most coaching programs cannot be evaluated, which is why they get cut.

Stop trying to measure coaching with revenue. Quota attainment, win rate, and pipeline coverage move for a dozen reasons at once — territory changes, pricing, competitive shifts, a good quarter in one vertical. Attributing a movement in any of them to coaching is a story, not a measurement. It also takes two quarters to see, by which point the program has already been judged.

Measure the behavior you coached, then the conversion step it should move. In that order.

  1. Pick one behavior.

    One. Not a program of six competencies. Discovery question count, or next-step confirmation, or talk ratio.

  2. Define the threshold in advance, in a number.

    "More discovery" is not measurable. "Eleven or more questions in a first meeting" is.

  3. Record where the team sits now.

    Before you coach anything. This is the step most often skipped and the one that makes everything afterward interpretable.

  4. Coach that one behavior for a defined window.

    Six to eight weeks.

  5. Measure the behavior again.

    Did it move? This is the honest test of whether coaching worked, and it's available long before revenue tells you anything.

  6. Then check the conversion step downstream.

    If discovery improved and first-meeting-to-opportunity conversion didn't, you learned something important: that behavior wasn't the constraint. That's a useful result, not a failed program.

The variance test. There's a diagnostic worth running before any of this, because it tells you whether coaching is your constraint at all. Cut your first-meeting-to-qualified-opportunity conversion by lead source, then by individual seller with the source held constant.

If conversion varies mostly by source, you have a sourcing problem and coaching won't fix it. If the same sources produce different outcomes depending on who takes the meeting, that variance is execution — and execution is coachable. I've written about that test in more detail in Pre-Call Intelligence: What It Is, What It Costs, and How to Implement It.

Three things worth tracking over time:

  • Behavior frequency against a defined threshold. The percentage of conversations meeting the bar, not the average.
  • Variance across the team. This is the real signal that a standard is holding. A narrowing spread means the standard is being applied consistently. A wide spread means each manager is still coaching their own.
  • Time to productivity for new hires. The clearest evidence that development has become systematic rather than personal.

One thing not to track: coaching hours logged. It measures activity, not development, and once it becomes a reported metric it starts being satisfied rather than achieved.

How to build a sales coaching program that scales

There are five real ways to address this, and I want to give each one a fair account before telling you what I'd do — including the two that involve buying nothing.

1Hire better sellers.
What it gives you

It works. Experienced sellers arrive with judgment and need less development.

What it costs

It doesn't compound, it gets more expensive as you grow, it doesn't survive a hiring market turning against you, and it makes the organization's capability a function of recruiting rather than a function of anything you control. Most organizations that try to hire their way out of this end up with a wider distribution, not a better one.

2Train more.
What it gives you

Genuinely valuable at onboarding and at a methodology change, when there's a body of knowledge to transfer.

What it costs

The limitation is well documented and every leader has experienced it: training fades without reinforcement, and reinforcement is coaching. Training is a good answer to "they don't know," and a poor answer to "they know and don't do it."

3Build your own rubric from your own best conversations.
What it gives you

Free, fast, and better than nothing, which is where most organizations start, correctly.

What it costs

Two costs. It inherits your ceiling, for the reasons in the section above. And it takes real senior time to build and maintain, which is why most home-built rubrics are written once and never revised.

4Buy conversation intelligence.
What it gives you

Solves coverage properly, which was a genuine problem. Gives managers visibility they could not otherwise have and removes sampling bias from review.

What it costs

The honest limitation is the one I've described: the benchmark remains internal. You'll know what happened across every conversation and you'll be comparing it to your own distribution.

5Adopt a standard sourced from outside your organization.
What it gives you

Solves provenance, which is the thing nothing else on this list solves.

What it costs

An external standard is not automatically right for you. It needs adapting to your deal types, your sales motion, and your buyer. A benchmark drawn from first conversations in enterprise technology is a poor fit for late-stage renewal negotiations in another category, and anyone who tells you otherwise is selling.

What I'd actually recommend, for most organizations.

Not one of these. A combination, in a specific order.

Start with the diagnostic in the previous section, before buying anything, so you know whether coaching is your constraint. If it is, get coverage — you cannot coach what you cannot see, and manual review will not get you there.

Then get a standard with provenance from outside your own distribution, and adapt it deliberately to your motion rather than adopting it wholesale. Then keep managers in the loop, because judgment about a specific seller in a specific situation is not something any system supplies.

The mistake I see most often is buying one of the five and considering the problem addressed. Coverage without a standard produces well-measured mediocrity. A standard without coverage produces a document nobody applies. Managers without either produce the memory-based coaching we started with.

Where we fit, and what we sell

I should be straightforward about my position, because it's a conflicted one and you should weigh what I've written accordingly. I've just argued that standards need external provenance, and we sell a product built on an external standard. That argument is commercially convenient for me and you should read it with that in mind.

Evolve Coach analyzes a conversation after it happens and turns it into development for the seller and visibility for the manager — scored against the benchmarks in our corpus rather than against a rubric we invented for the product.

The distinction I care about most is that Coach develops rather than grades. A score that tells a seller they were below the bar and nothing else is an evaluation, and evaluation without a path forward is waste. The purpose of capturing what happened in a conversation is growth

We didn't build a model and go looking for a market. We ran these conversations for six and a half years, scored them because we were accountable for whether the pipeline converted, watched what separated the ones that became revenue from the ones that didn't, and built the product around that. The AI is how it's delivered. It isn't the reason it works.

The chain we believe in, stated link by link rather than compressed: a standard with provenance produces consistent coaching; consistent coaching produces changed behavior; changed behavior produces better conversations; better conversations convert to qualified opportunities more often. Each link is testable. None of them is a promise about pipeline.

If you run the diagnostic and find your variance sits with your lead sources rather than your sellers, coaching isn't your constraint — and you should go and work on the thing that is. I'd rather tell you that than sell you this.

Convert more leads you already pay for

Coach scores every conversation your sellers run against 35,000 real buyer conversations from 1,200+ IT vendors, then turns each one into a personalized action plan for the seller who ran it. Be among the first to use it.

Join the waitlist

FAQ

What is sales coaching?

Sales coaching is the ongoing development of a seller's ability to run better customer conversations, delivered through specific, repeated feedback anchored in real conversations. It differs from training, which transfers knowledge once, and from pipeline review, which inspects deals rather than developing capability.

How do you measure sales coaching effectiveness?

Measure the behavior you coached before the revenue you hoped for. Pick one behavior, define its threshold as a number, record where the team sits before coaching, coach for six to eight weeks, then measure the behavior again. Only after it moves should you check whether the conversion step downstream moved with it.

How do you measure the ROI of sales coaching?

Isolate a single behavior and a single conversion step, then compare the conversion rate for conversations meeting the behavioral threshold against those that don't. That gives a defensible link between behavior and outcome. Attributing quota attainment directly to coaching is not a measurement, because quota moves for many reasons at once.

Why doesn't sales coaching scale?

Three structural reasons: it runs on whichever conversations a manager happened to observe, it varies with whichever manager a seller reports to, and it is rarely tied to a written standard, so nobody can tell afterward whether it worked. More manager time addresses none of the three.

Why is sales coaching inconsistent across a team?

Because the standard usually lives in individual managers rather than in the organization. Each manager coaches toward what they personally value or were once good at. Consistency requires a written, numeric standard that exists independently of who is doing the coaching.

How often should sales coaching happen?

Frequently enough that feedback still connects to a specific conversation the seller remembers — weekly or fortnightly for most teams. Cadence matters less than specificity. A short session tied to one real conversation and one behavior outperforms a long monthly review of general impressions.

What should a sales coaching session include?

One specific conversation, one behavior, an objective observation of what happened, and a written commitment the seller owns going into the next conversation. Sessions that cover several competencies at once tend to change none of them.

How do you write a sales coaching plan?

Define the behaviors you'll coach and the numeric threshold for each, record the team's current position against them, assign one behavior per seller per cycle, set the review cadence, and decide in advance what evidence would show the plan worked. A plan without a baseline cannot be evaluated.

How can AI support sales coaching?

Its clearest contribution is coverage: analyzing every conversation rather than the small sample a manager can review, which removes sampling bias. What it does not do is validate the standard being applied. Scoring every conversation against an unevidenced rubric produces more measurement, not better coaching.

What should you look for in sales coaching software?

Coverage, integration with how your team already works, and — the question most evaluations skip — where the platform's benchmarks come from. Ask directly whether the standard is derived from your own conversations or from evidence outside your organization, because that determines whether you're measuring against good or against your current ceiling.

What's the difference between sales coaching and sales training?

Training transfers knowledge and is finite: onboarding, a product launch, a methodology rollout. Coaching develops capability and is continuous, working on one behavior at a time against real conversations. Training answers "they don't know." Coaching answers "they know and don't do it."

What does good sales coaching actually look like?

Specific rather than general, anchored in a real conversation rather than an impression, focused on one behavior at a time, measured against a written threshold, and consistent regardless of which manager delivers it. If two managers on the same team would coach the same conversation differently, the organization doesn't yet have a coaching system.