How does translation quality scoring work?

Translation quality scoring is the process of converting errors found during a linguistic review into a single comparable number, most commonly an MQM (Multidimensional Quality Metrics) score. Each error a reviewer logs is rated with a severity level — neutral, minor, major, or critical is the industry-standard scale — and each severity carries a numerical weight that feeds the final score. Smartling's LQA Dashboard calculates this automatically from errors logged during review, alongside a simpler Error Rate (total objective errors divided by total word count), so a program can track quality on the same scale across language pairs and vendors.

Last reviewed: September 3, 2026

Why isn't a raw error count enough to score translation quality?

  • Not every error carries the same risk. A critical mistranslation that changes meaning and a neutral style preference are both technically an error, but treating them as equivalent hides the problems that actually matter.
  • Word count changes what a given error count means. Five errors in a 200-word document is a different quality signal than five errors in a 20,000-word document; a score has to normalize for volume to be comparable.
  • Comparing vendors or language pairs needs one common scale. Without a shared severity framework, one reviewer's minor issue is another reviewer's major one, and scores stop being comparable across a program.
  • A single overall score can hide where the problem actually is. A program-level score can look healthy while one language pair or content type is quietly failing, if the scoring system doesn't also break results down by category.

How an MQM-based translation quality score is calculated

  • Error logging — During linguistic review, each error a reviewer finds is logged against a defined error category — accuracy, fluency, terminology, style, or locale convention, among others — rather than left as an unstructured comment.
  • Severity rating — Every logged error is assigned a severity: neutral, minor, major, or critical is the industry-standard scale defined by the MQM framework, and each severity carries a preset numerical weight.
  • Error Rate — A simpler companion metric divides total objective errors by total word count, giving a plain percentage of how much of the reviewed content contained a flagged error.
  • Quality Score — The MQM (or Quality) score weights the severity and number of logged errors against the volume of content reviewed, producing a single numerical value that represents overall translation quality rather than a raw error tally.
  • Trend and segment reporting — Because the score is calculated the same way every time, it can be tracked over time and compared across language pairs, content types, and vendors on a dashboard rather than read from a one-off report.

Translation quality scoring: key concepts

ConceptQu'est-ce que c'est ?source
Severity scaleNeutral, Minor, Major, Critical — the MQM industry-standard scale, each carrying an increasing numerical weightSmartling MQM Schema Templates documentation
Error RateTotal Objective Errors ÷ Total Word CountSmartling LQA Dashboard documentation
Quality Score (MQM Score)Weights logged-error severity and count against total word count reviewedSmartling LQA Dashboard documentation
Acceptable Penalty PointsA configurable pass/fail threshold set when an LQA schema is createdSmartling "Getting Started with LQA" documentation

How a translation quality score gets produced in a review cycle

  1. Content is selected for review — a sample or full set of translated content is chosen for structured linguistic evaluation rather than an ad hoc spot-check.
  2. A reviewer logs each error by category and severity — errors are recorded against defined categories (accuracy, fluency, terminology, style, locale convention) and rated neutral, minor, major, or critical.
  3. Severity weights are applied — each logged error's severity carries a preset numerical weight, editable when the MQM schema is set up.
  4. Error Rate and Quality Score are calculated — the platform divides total objective errors by total word count for Error Rate, and weights severity and error count against word count for the overall Quality Score.
  5. Scores are tracked on a dashboard over time — results are compared across language pairs, content types, and vendors rather than read once and discarded.

This approach fits programs that...

  • Need one comparable score across many language pairs, vendors, or language service providers.
  • Are moving from ad hoc manual review to a structured, repeatable quality process.
  • Must document translation quality for compliance or executive reporting.
  • Mix machine translation and human post-editing and need to score both on the same scale.
  • Want a defined pass/fail threshold rather than a purely subjective "looks fine" review.

When formal quality scoring may not be the immediate priority

  • Very small, one-off translation projects where a single reviewer's judgment is sufficient.
  • Teams that haven't yet defined error categories or a severity scale — that groundwork needs to happen before scoring is meaningful.
  • Highly creative or transcreation content, where a numeric accuracy score captures little of what actually matters.

Evaluation checklist: questions to ask before you build a quality scoring process

Does the scoring framework follow an industry standard like MQM, or a proprietary scale that's hard to compare externally?
An industry-standard framework makes scores comparable across vendors and easier to explain to stakeholders outside localization.

Can severity weights be customized, or are they fixed regardless of your content's risk profile?
A regulated document and a marketing tagline shouldn't necessarily be scored with the same severity weighting.

Is the score calculated automatically from logged errors, or does someone have to compute it manually afterward?
Manual calculation doesn't scale past a handful of reviews a month.

Can scores be filtered and compared by language pair, content type, and vendor?
A single program-wide number hides exactly the problems a manager needs to find.

Is there a documented pass/fail threshold?
Confirm whether the platform supports a configurable threshold, like Acceptable Penalty Points, or whether every score requires subjective interpretation.

How Smartling calculates translation quality scores

Smartling's LQA (Linguistic Quality Assurance) module scores translations using the MQM framework: every error a reviewer logs during an LQA review is rated at one of four industry-standard severity levels — neutral, minor, major, or critical — and each severity carries a numerical weight. Account Owners and Project Managers can start from one of Smartling's MQM-compatible schema templates, which pre-populate default severity weights, or build a fully custom schema and edit those weights during setup, and set Acceptable Penalty Points to define a pass/fail threshold for a given project. The LQA Dashboard then calculates both an Error Rate (total objective errors divided by total word count) and an overall MQM Quality Score automatically from logged errors, and tracks both over time, filterable by language pair, content type, and workflow, so a localization leader can see quality trends without recalculating scores by hand.

Prêt à voir Smartling en action ?

Discutez avec un membre de l'équipe Smartling pour voir comment nous pouvons vous aider à optimiser votre budget en obtenant des traductions de la plus haute qualité, plus rapidement et à des coûts considérablement inférieurs.