Skip to content
Notes

Short read · Building

A two-card pilot picked the wrong AI translator

On October 3 I chose a translation model twice, for my plant encyclopedia and for this blog. The first pilot fooled me; here is how, and the blind tests that replaced it.

Max Kiriienko
Tech Lead SEO & Marketing · 8 min read

Читати українською

On October 3 my houseplant encyclopedia, plants.place, was getting an English version. 133 species cards — one page per plant — had to go from Ukrainian into English, and I needed a model to do it.

OpenAI’s pricing page lists models called sol, luna, astra and terra, but doesn’t say what the names mean. So my agent ran a pilot: two tricky cards through four models. One was the snake plant, with Ukrainian folk names to handle. The other was a dendrobium orchid, where shops sell hybrids under the species name.

The pilot: two cards, four models

ModelPrice per 1M tokens, in / outAll 133 cardsPilot verdictBlind test
gpt-6-luna$0.10 / $0.50≈ $0.22Best: the most natural EnglishLost: won 2 of 10
gpt-6.1-sol$2 / $10≈ $4.17Accurate, wooden in placesWon 8 of 10
gpt-6-astra$10 / $50≈ $20.84Accurate and cleanBecame the judge
gpt-5.5$5 / $30≈ $12.34Worst: Cyrillic left in the English textDropped

Prices are OpenAI’s list prices on October 3, 2026. The cost of all 133 cards is the pilot’s estimate.

All eight translations passed the automatic checks: the JSON structure, every number and every Latin name were intact. The cheapest model read best. On the snake plant it explained what the Ukrainian folk names mean — “mother-in-law’s tongue” and “pike’s tail” — and then gave the English ones. At $0.22 for the whole English version, it looked like an easy call.

The pilot report itself said two cards were too few. So before the full run there was one more step: a blind test.

The blind test reversed it

Ten cards, two candidates: gpt-6-luna and gpt-6.1-sol. The judge was gpt-6-astra, the most expensive of the four. It saw the Ukrainian source and two translations marked A and B, in a new random order for every card. It scored four things: fidelity (nothing added, lost or changed), natural English, correct gardening terms, and the handling of Ukrainian folk names. It had to quote every problem it found.

gpt-6.1-sol won 8 to 2, with an average of 8.9 against 8.2. The test took two and a half minutes and cost $1.50. Both pilot cards were among the ten, and the judge gave both to the model the pilot had ranked lower.

For the whole English version, the difference between the two models was about $4.

Why the pilot got it wrong

The judge’s notes explain it. The cheap model often read better. The other one was more faithful to the source on nine cards of ten.

What the blind judge saw on 10 plant cards

gpt-6.1-sol was more faithful to the source on 9 cards, gpt-6-luna on none, one tie. gpt-6-luna read more naturally on 5 cards, gpt-6.1-sol on 1, four ties.

Ties: one on fidelity, four on natural English. On eight cards the judge scored each criterion; on two it said it in words. Source: blind A/B test on 10 plants.place cards, judge gpt-6-astra, Oct 3, 2026

Here is what that looks like in one line of the snake plant card:

Ukrainian source

«Після підсихання» — after the mix has partly dried. Summer: once every 7 days as a guide.

gpt-6-luna

"After the potting mix dries out." "In summer, aim for once every 7 days." Reads well. Implies the mix should dry out fully, and turns a guide into a target.

gpt-6.1-sol

"After partial drying." "In summer, a guideline is once every 7 days." Stiff. Says what the source says.

Live card now

"Water after the mix has dried out a little." "As a guide, water once every 7 days in summer…"

The watering line of the snake plant card in the blind test, October 3. The live card comes from the later full run, with a longer prompt.

“After the potting mix dries out” sounds right. Taken literally, it tells a reader to wait for a dry pot, while the source says to water once the mix has partly dried. Across 133 care cards, slips like this add up to wrong advice in fluent English.

The same slip was already in the pilot: there, too, the cheap model wrote “After the soil dries out”. The pilot verdict didn’t mention it. When you read for an impression, you judge how a translation sounds. Fidelity shows only when someone puts it next to the source, line by line, with a list of what to check.

What the pilot was good for

It caught a real failure. gpt-5.5 left «Дендробіум нобіле» in Cyrillic in the middle of an English sentence. It also took the Ukrainian explanation of the folk names and pinned it on the English names. Leftover Cyrillic shows up in one card and is easy to check by code. “No Cyrillic in English text” became a rule in the validator the same day.

For the winning model, the pilot’s cost per card held up: about $0.031 in the pilot, about $0.035 in the full run. Its timing didn’t: 37 seconds per card in the pilot, often several minutes per card in the full run.

So a pilot is good at catching what’s broken. It is bad at ranking two good models, where the difference is in small shifts of meaning.

The second time, I went blind from the start

That evening I needed a translator for this blog, from English into Ukrainian. The default in my translation script was gpt-5.5. This time there was no pilot. My agent took ten paragraphs from the first drafts of my posts and translated them with three models and the same instructions. Three Claude agents judged them blind, each with a different brief: a copy editor, a bilingual SEO and a strict proofreader. The order of the three versions changed from paragraph to paragraph.

ModelPrice per 1M tokens, in / outAverage scoreBest versionJudges 1 / 2 / 3
gpt-5.6-sol$4 / $208.4017 of 308.50 / 8.15 / 8.55
gpt-5.5, the old default$5 / $307.7010 of 307.50 / 7.90 / 7.70
gpt-5.6-terra$2 / $127.233 of 307.00 / 7.50 / 7.20

Scores out of 10. Thirty verdicts: ten paragraphs, three judges. Prices as of October 3, 2026.

All three judges put gpt-5.6-sol first, so it became the default. The old default came second, and it was also the most expensive of the three.

Would a two-paragraph pilot have found the winner? I re-sliced the same scores: every possible pair of the ten paragraphs, with the same three judges.

If the test had been smaller: how often another model would have won

Out of every possible set of the ten paragraphs: with 1 paragraph another model wins 40% of the time, with 2 — 22%, with 3 — 13%, with 4 — 7%, with 5 — 4%, with 6 — 1%. From 7 paragraphs on, every set picks gpt-5.6-sol.

Every possible set of the ten paragraphs, scored by the same three judges. A miss: another model had the higher average, or a tie. This re-slices one test and assumes its full result is right. Source: blind test of three translation models, 10 paragraphs × 3 judges, Oct 3, 2026

10 of the 45 pairs pick another model. The paragraphs aren’t equal: all three judges gave paragraphs 4 and 5 to gpt-5.5, and paragraph 3 to gpt-5.6-terra. Start with those, and you keep the old default with a clear conscience. With six paragraphs, 3 of 210 sets still miss. From seven up, none do.

The winner still needed rules

Some mistakes showed up in more than one model. Two of the three translated “SERP feature” as «функція» — the Ukrainian word for a software function. Ukrainian SEO people say «елемент видачі». For “intent”, one model left the English word, one wrote «намір» and one «інтенція». None of them, not even the winner, wrote «інтент», the word practitioners use. gpt-5.6-terra also moved a number: “compressed 43 times” became “exceeded the agent’s memory 43 times”.

So the most useful output of the test wasn’t the winner. It was 16 word-level rules, most of them straight from the judges’ notes, plus a general one: no English words left in Latin script. They went into the translation prompt the same evening.

The translation found errors in the original

While the full run on plants.place was going, the same judge checked ten other cards that were already translated: 7.9 out of 10 on average, for $1.08. It flagged six critical problems. Four of them traced back to the Ukrainian original, not to the translation. The worst: our Ukrainian card said to winter the Venus flytrap in shade, only 4–5 °C cooler than in summer. A flytrap needs a cool rest in bright light. Both versions now say so.

The two critical errors that were the translation’s own were about meaning, not grammar. One was an extra English name for the mandevilla that, according to the judge, isn’t used for this plant. The other: a question about cats changed direction. “Can I keep it at home if I have a cat?” became “Is it safe to keep at home if I have a cat?”, so the answer “Yes, the plant is toxic” turned into a contradiction. The Ukrainian answer was clumsy too, and we fixed it in 12 cards. A less serious slip came up in three cards: «перегній» in a potting mix became “humus”, while here it means well-rotted manure. The cat question and the manure are now rules in the prompt.

What I can’t claim

  • The judges are models too. On plants.place it was one judge from the same company as the two contestants. On the blog it was three judges from another company.
  • Ten samples is a small test as well, just less small. And 8 to 2 is really 7 clear wins, 1 clear loss and 2 equal scores where the judge still named a winner — one each way.
  • The random order put the winner first on 8 of the 10 cards. It also won both cards where it came second, and both of the other model’s wins came from the second slot. I see no position effect, but I can’t rule one out. On the blog, two of the three judges saw the versions in the same order.
  • While writing this note, two of the judge’s quotes were checked against the source: the watering line, and “bare stems” for «голі пагони», which here means hairless stems. Both hold. Nobody has read all twenty translations line by line.
  • Model names and prices change fast. Everything here is as of October 3, 2026.

How I choose an AI translation model now

  1. Run a pilot to catch what’s broken: lost numbers, a broken structure, the source language left in. Don’t rank models on it.
  2. Rank on ten or more samples, blind. Shuffle the order for every sample.
  3. Give the judge the source and written criteria. Fidelity first, natural language second. Ask it to quote every problem.
  4. Use a judge from another model family, or several judges with different briefs.
  5. Check the hard things in code: numbers, names, no leftover script.
  6. Turn repeated mistakes into rules in the prompt. The winner makes them too.
  7. Check a few of the judge’s quotes against the source yourself.