Skip to content
Compliance included · one account for numbers, voice and SMS · 191 countries covered.Compliance included · one account for numbers, voice and SMS · 191 countries covered.Compliance included · one account for numbers, voice and SMS · 191 countries covered.Compliance included · one account for numbers, voice and SMS · 191 countries covered.
Twiching
All posts

What Actually Determines Wholesale VoIP Voice Quality

T
Author: Twiching TeamVoice Quality Engineer
August 22, 20269 min read
What Actually Determines Wholesale VoIP Voice Quality

Introduction

A route can post a 65% ASR, five-nines uptime, and still sound bad. Answer rate and availability tell you whether a call connects and stays up — they say almost nothing about whether the two people on that call can actually understand each other clearly.

Voice quality is a separate axis entirely, governed by a different set of factors most buyer's guides skip past on their way to rate decks and SLA percentages: how the audio gets encoded, how echo gets suppressed, how the receiving end smooths over network hiccups before they become audible. This guide is about that second axis — the actual acoustic experience of a wholesale VoIP call — covering MOS scoring, codec tradeoffs, echo cancellation, jitter buffer behavior, and packet loss concealment, along with how to test for real audio quality instead of inferring it from network statistics alone.

Why Connection Metrics Don't Measure How Calls Sound

Why Connection Metrics Don't Measure How Calls Sound

ASR, ACD, and uptime are connection-layer metrics — they describe whether a call happens and how long it lasts, not what it sounds like while it's happening. A route can deliver a perfect connection record while still producing calls that sound muffled, echo-laden, or intermittently choppy, because none of those symptoms affect whether the call connected or how long it stayed up.

This gap matters because a buyer evaluating routes purely on connection statistics can end up with a technically reliable route that customers still complain about — the complaints just show up as "call quality issues" in support tickets rather than anywhere on a route-performance dashboard.

MOS: The Score Built Specifically to Measure This

Mean Opinion Score (MOS) is a standardized 1-to-5 rating of perceived audio quality, originally derived from human listeners rating sample calls and now commonly estimated algorithmically in real time for production traffic. A MOS above 4.0 is generally considered toll-quality — indistinguishable from a traditional phone call to most listeners. Scores in the 3.5–4.0 range are noticeably degraded but usually still usable; below 3.5, quality issues become distracting enough to affect the conversation itself.

  • 4.0–5.0 — toll-quality, comparable to a traditional landline call
  • 3.5–4.0 — noticeably degraded but generally functional for normal conversation
  • Below 3.5 — quality issues become distracting; below roughly 3.0, calls become genuinely difficult to conduct

A provider that can report real-time or near-real-time MOS per route, rather than only ASR and PDD, is measuring something meaningfully different from — and arguably more directly relevant to — the actual customer experience than connection statistics alone.

Codec Choice and Its Direct Effect on Clarity

Codec Choice and Its Direct Effect on Clarity

The codec a call uses sets a ceiling on how good it can possibly sound, independent of everything else in the network path. G.711 encodes audio with minimal compression, delivering the closest thing to toll-quality sound at the cost of higher bandwidth per call. G.729 compresses more aggressively, trading some clarity for significantly lower bandwidth use — noticeable on close listening but often acceptable for routine business calls. Opus, increasingly common on modern wholesale routes, adapts its bitrate dynamically to network conditions and can match or exceed G.711 quality while using less bandwidth in good network conditions.

Codec mismatches between what your system offers and what a route actually negotiates can silently downgrade quality below what either side's headline specs suggest — confirming which codec a call actually used, not just which ones are theoretically supported, is worth verifying rather than assuming.

Echo Cancellation and Why Some Routes Sound Hollow

Echo on a VoIP call comes from acoustic or electrical reflection of the sender's own voice back to them, delayed just enough to be perceptible and distracting. Echo cancellation algorithms model and subtract that reflected signal in real time, and how well a route implements this directly affects whether calls sound clean or subtly (or not so subtly) hollow.

Echo problems are often route-dependent rather than universal — a provider's network might handle echo cleanly on most paths but poorly on specific interconnects, particularly ones involving older TDM-to-IP gateway conversions where analog echo is more likely to enter the signal path in the first place. A route that sounds fine on a quick test call can still have echo issues that only surface on longer conversations or specific destination types, which is part of why quality testing benefits from more than a handful of short calls.

Jitter Buffers and Packet Loss Concealment

Jitter Buffers and Packet Loss Concealment

These two mechanisms work together to hide network imperfections from the listener, and how well they're tuned determines whether minor network hiccups are inaudible or obviously disruptive. A jitter buffer holds incoming audio packets briefly and releases them at a steady pace, smoothing out the uneven arrival timing that real networks produce — a buffer that's too small lets jitter through as choppy audio, while one that's too large adds noticeable conversational delay.

Packet loss concealment (PLC) handles the packets that never arrive at all, using algorithms to intelligently estimate and fill the gap rather than leaving silence or an audible glitch. Good PLC can make occasional packet loss essentially imperceptible; poor implementations turn the same loss rate into an obviously broken-sounding call. Two routes with identical measured packet loss percentages can sound completely different depending on how well each one's PLC is implemented — another reason raw network statistics alone don't fully predict perceived quality.

How to Actually Test Voice Quality Before Committing

How to Actually Test Voice Quality Before Committing

Testing for voice quality specifically requires listening, not just monitoring dashboards — network statistics can look clean while the audio still has problems the numbers don't capture.

  • Place real conversational-length calls, not quick connection tests — echo and jitter-buffer problems often surface only after a minute or more of continuous speech
  • Test across multiple destinations and times of day — quality can vary by route and by network congestion patterns that a single test window won't reveal
  • Listen specifically for echo, choppiness, and unnatural pauses — these are the audible symptoms of the mechanisms covered above, not abstractions on a report
  • Ask for real-time MOS data per route if the provider offers it, and compare it against your own subjective listening rather than trusting either signal alone

A provider confident in their voice quality will support this kind of testing without resistance — reluctance to let you actually listen to real call quality before committing is itself a signal worth taking seriously.

Conclusion

Voice quality and connection reliability are related but genuinely separate problems, measured by different signals and solved by different engineering. ASR and uptime tell you whether calls happen; MOS, codec choice, echo cancellation, and jitter buffer tuning tell you whether those calls are actually pleasant to be on.

A wholesale VoIP evaluation that stops at connection metrics is only answering half the question a customer actually cares about — the other half only shows up when you listen to a real, full-length call and judge it the way an actual caller would.

FAQ

Questions about Twiching, answered.

A MOS above 4.0 is generally considered toll-quality, comparable to a traditional phone call. Scores between 3.5 and 4.0 are noticeably degraded but usually still functional, while scores below 3.5 indicate quality issues significant enough to affect normal conversation.

You reached the end

Read the next one. Slide to continue.

What is Wholesale Voice? Powering Scalable and Cost-Effective Communication

Drag the dial fully to the right to open the next dispatch.

Get started

One workspace, every conversation.
See what a real phone stack does.

Phone numbers, voice, SMS and AI on one account. No credit card required to get started.

Compliance with applicable regulations required.