How to Run a Translation Vendor Bakeoff

    #AI#document#translation#BluTranslate#Bluente#enterprise#comparison#security#financial#compliance#authenticity#localization#format#preservation

    A vendor bakeoff is a controlled head-to-head test: the same documents, the same term base, the same deadline, given to every candidate, with outputs stripped of identifying marks and scored against a rubric agreed before anyone sees a result. It works when the document set includes your genuinely difficult files — the scanned annex, the merged-cell table, the deck with text in the master — and fails when someone submits clean samples, because every vendor passes on clean samples and you learn nothing. Half the scoring needs no linguist at all: numbering, cross-references, table structure, numbers and terminology can be checked mechanically.

    This is a protocol you can run in about three weeks: how to choose the files, how to blind them, who reviews what, how to score, and how to write the decision down.

    Assembling the Document Set

    Pull the set from real work, not from a sample folder. Five to eight documents per language pair is enough if they are chosen for coverage rather than convenience.

    Include, deliberately: a long agreement with multi-level clause numbering and internal cross-references; a rate table with merged cells, footnotes and a totals row; a scanned PDF; a deck where text lives in a master, a chart label and a grouped shape; a spreadsheet with formulas and named ranges; and one document that is genuinely badly made — manual numbering, tracked changes left in, a table built out of tab stops.

    That last one is not a trick. Someone in your organisation produced it, and someone will send it to the vendor in month two. If a system only works on well-formed input, you need to find out in the bakeoff rather than in production.

    Cap the total volume so every vendor can turn it around inside the same window without special handling. If the eventual use is diligence, draw from the same material you would send to a due diligence review.

    Freezing the Test Conditions

    The comparison is only valid if the only variable is the vendor. That takes more discipline than it sounds.

    Send identical source files — the same bytes, not a re-exported copy. Send the same term base to every vendor, in the same format, and ask each to confirm it loaded; TBX under ISO 30042 is the interchange format to ask for, because a vendor that can only accept a spreadsheet will struggle to accept anything else later.

    Give identical instructions, including target locale, treatment of do-not-translate items, and what to do with untranslatable content. Give the same deadline, and record actual turnaround against it.

    Say nothing about how the output will be scored beyond the rubric you share with everyone. Do not tell a vendor which document you consider hard. And take the files back in their original formats — a vendor that returns a PDF when you sent a Word file has already answered a question about their pipeline.

    Blinding the Outputs

    Reviewers score vendors they recognise differently, without intending to. Incumbents get benefit of the doubt; the cheap challenger gets scrutiny; the brand everyone has heard of gets a halo. Blinding is cheap insurance.

    Before review, one coordinator — not a reviewer — strips document properties and metadata, renames files to neutral codes, and removes headers, footers, watermarks or cover pages that identify the source. Keep the mapping in a single file that reviewers cannot see. Randomise the code assignment so vendor A is not always sample 1.

    Blinding survives contact with reality imperfectly: house style, a distinctive way of handling untranslated strings, or a particular font substitution can give a system away. Accept partial blinding rather than skipping it. What matters is that the reviewer is not reading a brand name at the top of page one when they decide whether clause 9 is acceptable.

    Unblind only after every score is recorded and locked.

    The Mechanical Checks That Need No Linguist

    Run these first, on every document, before a reviewer opens anything. They are objective, they are fast, and they eliminate candidates without spending expensive attention.

    Clause and list numbering. Extract the numbering sequence from source and target and compare. Restarts, skipped levels and lists converted to literal text are all disqualifying defects on a structured agreement.

    Cross-reference resolution. Every "see clause 12.4" must point at the clause that is now 12.4. Word field codes either survive the round trip or silently become static text.

    Table structure. Row and column counts per table, merged-cell topology, whether totals rows are still aligned with what they total.

    Numeric and date integrity. Extract every figure, percentage, currency amount and date from both versions and diff them. Watch decimal separators and day-month ordering in particular.

    Terminology enforcement. Count occurrences of each approved term and list every deviation, with location — this is what glossary enforcement has to deliver in practice.

    Completeness. Untranslated segments, dropped footnotes, text left in charts, images and text boxes.

    The Reviewer Rubric

    For the linguistic half, use an error typology rather than an impression score. The MQM framework gives you categories — accuracy, terminology, fluency, style, locale convention — and a severity scale, which means two reviewers can disagree productively instead of trading adjectives.

    Choose reviewers who know the subject and the target jurisdiction, not just the language. A native speaker who has never read a facility agreement will flag register and miss an inverted obligation.

    Two reviewers per language pair, with a shared calibration passage scored before the real pass, is the minimum that tells you anything about your own measurement reliability. Give them a fixed error inventory and a severity definition written in your terms: what counts as critical for your documents, in one sentence each.

    Cap the reading. Ask for a defined word count per vendor rather than "read it all", or the last vendor reviewed gets the least attention.

    Scoring and Weighting

    Agree the weights before the results exist. Afterwards, every weighting argument is really an argument about which vendor should win.

    A workable structure: mechanical checks form a pass/fail gate plus a defect count; linguistic review produces per-dimension penalty scores normalised per thousand words; operational factors — turnaround against deadline, format fidelity of the returned file, responsiveness, whether the term base actually loaded — carry their own weight.

    Two rules keep it honest. Critical errors are counted separately and never averaged away; one inverted obligation is not offset by fluent prose elsewhere. And the reformatting effort each output requires gets an estimate in hours, because that is a cost the sticker price does not include, and on document work it is frequently the largest difference between vendors.

    Record everything at document level, not just in aggregate. The pattern of which vendor failed which document is the finding.

    The Easy-File Trap

    The most common way a bakeoff goes wrong is that it is run on clean, short, well-structured text. Every serious vendor scores well, the scores cluster within noise, and the decision defaults to price — at which point the whole exercise was theatre.

    It happens for understandable reasons. Confidentiality makes people reach for sanitised samples. Whoever assembles the set picks documents that are easy to explain. Vendors, reasonably, prefer clean input. Nobody wants the test to embarrass the files.

    The corrective is a rule: at least half the set must be material that has previously caused a problem. Ask the people who actually handle these documents which files they dread, and use those. Threads like this one on translating a 100-page Word document without wrecking the formatting are a decent proxy for the failure modes worth testing — long, structured, styled documents are where systems separate.

    If a vendor objects that the test files are unrepresentative, note the objection and continue. They are representative of you.

    Recording the Decision

    Write a short memo while the detail is fresh: the document set and why each file was included, the conditions, the rubric and weights with their date of agreement, the per-vendor results at document level, the unblinding key, and the decision with its reasoning.

    This does two jobs. It makes the choice reviewable — by a procurement committee now, by a regulator or an internal auditor later — and it gives you a fixed baseline. Re-run the same set at renewal and you can say whether the incumbent improved, which no amount of relationship management will tell you.

    Keep the sample set under change control alongside your term base and whatever certification evidence you rely on. Bluente is straightforward to include in a bakeoff for exactly this reason: the structural checks above are the ones the pipeline is built to pass, across 120+ languages and 22+ file types, and structural results are the ones that survive an argument.

    Sources and Further Reading

    Related Reading

    Last reviewed 24 August 2026 by the Bluente document engineering team, who build and test the pipeline described here. We update these guides when the underlying standards, regulations or file formats change.


    Test the files you dread, not the files you'd demo. Try BluTranslate free.

    Published by
    #AI#document#translation#BluTranslate#Bluente#enterprise#comparison#security#financial#compliance#authenticity#localization#format#preservation
    Back to Blog
    Share this post: TwitterLinkedIn