In localization, source file context has always mattered, but the numbers show that even now, many teams still deliver it through channels a linguist never actually sees while working. This article discusses the results of a test to determine whether a more direct approach made a difference.

The test consisted of generating a detailed translation brief with AI for a piece of mental health content, breaking it into segment-level annotations, and uploading those annotations straight into Localazy's UI. After AI translation, one human reviewer* was given zero context while another was able to see suggestions sitting right in the segment instead of buried in a separate document.

The tests we ran showed that providing accessible context resulted in improvements in translation quality, ranging from glossary consistency and dialectal accuracy to how much room a linguist feels they have for transcreation.

*Note: Two professional human reviewers were used for this test. Both were accredited linguists from Localazy's Human translation services team. The test was carried out inside the platform, with only one of them being able to see the real-time suggestions inside the UI.

💬 Where do teams get their context? 🔗

But let's start at the very beggining. Following up on my last article on AI in pre-translation workflows, I've kept digging into how other localization professionals actually handle source file analysis in practice. When I gave a talk on AI in pre-translation workflows at a recent localization industry event, I polled the room about something I'd long pondered: where were teams actually giving linguists the context they needed to do their jobs well?

Here’s what they said:

  • 📩 More than half of the respondents deliver task-specific instructions through project management tools, email, or similar channels outside the TMS entirely.
  • 📝 One quarter work directly in a CAT tool or source file.
  • 📘 Less than a quarter of respondents rely only on a style guide or glossary.
  • 👩🏻‍💻 One used ad-hoc chat.
  • ❌ And zero said they assigned tasks with no instruction at all.
How are task-specific instructions usually provided to translators? Pie chart with results.
Source: Data collected by Abby Frackenpohl at a localization event.

All of this tells us something: people know that context matters. But the channels they’re using aren’t easily visible at that exact moment a translator is staring at a segment that needs a judgment call.

People know that context matters, but often there isn't enough visibility for it. Over 52% of respondents say they use tools outside a TMS to provide it

When I asked which pre-translation task would deliver the most value if automated, glossary and terminology extraction came out on top with almost half the votes. Source QA followed and translator FAQ generation came in after that.

Cultural risk and sensitivity flagging (arguably one of the highest-impact interventions) got only 2 votes, which, in my opinion, says less about its value and more about how rarely PMs have seen it done well enough to truly move the needle. Independent of industry, when cultural nuance is made visible and suggestions are brought to the forefront, the right type of localization can make or break a product.

Which pre-translation task would be most valuable to automate with AI? Pie with results.

What made both results more interesting was that the features we were discussing (style guides, glossary insertion, contextual notes...) are precisely what most TMS platforms say they offer. The misalignment, however, is whether those features are populated with enough visibility to actually be useful, and whether that information reaches the linguist in the right place at the right time.

That's the problem I went back to Localazy to try to solve.

🕵️‍♀️ My test with integrated context suggestions 🔗

Most TMSs offer glossaries, style guides and translation memories, but I wanted to take it one step further and bring the context even closer to the linguist.

Step 1: Build the brief 🔗

The first step was creating something worth uploading. I used an excerpt from an article about burnout and resilience. It was an editorial piece in the wellbeing area of about 500 words long.

I gave Claude a detailed prompt covering the full picture of what a linguist working on this content would need to know: tone, register, target audience, the brand's particular pain points from previous localization cycles, areas where we'd want genuine transcreation rather than direct translation... I also flagged to watch for idioms and American cultural references that could undercut the goal of sounding local and authentic, and added a note to stay alert to clinical language that needed careful handling. To get an idea, you can find the full brief here.

A snippet of the brief I created using Claude with tone of voice remarks for translation.
A snippet of the brief I created using Claude with tone of voice remarks for translation.

Claude returned a comprehensive brief. I edited it down a bit (let’s be honest, no one will read a 10-page style guide) and uploaded it to Localazy to sit alongside the style guide that I had already created in the platform, where I included it as General instructions.

The instructions added on Localazy's Style Guide.
The instructions added on Localazy's Style Guide.

Step 2: Generate a JSON file 🔗

Then I took it a step further. I asked Claude to generate a JSON file annotating its specific suggestions against the segments where they applied. It flagged the specific strings where a term might need a glossary check and and a cultural reference appeared, along with more.

Step 3: Upload suggestions directly to Localazy 🔗

I uploaded that JSON file to Localazy, and here's where the concept started to become something tangible: the comments from the JSON mapped directly into the UI notes where linguists actually work.

They were visible right there, in the string!

The JSON comment (bottom left) was included directly in the appropriate string for context.
The JSON comment (bottom left) was included directly in the appropriate string for context. Neat!

💡 What the test revealed 🔗

To do the first translation pass, I chose Spanish (ESLA) as the target language for our mental health content and one result stood out immediately. The source content included a reference to a Band-Aid. 🩹 In ESLA Spanish, the culturally accurate equivalent isn't tirita (that's the term used in Castilian Spanish) and while Spanish-speaking audiences across Latin America will understand it, it doesn't read as local. In most of Latin America, the colloquial term is curita. Claude caught this and so the JSON I uploaded flagged the Band-Aid reference right in the CAT tool and noted the regional distinction.

On this first round, Localazy's AI suggestions, Google Translate, and DeepL all proposed tirita. On the second, thanks to the additional context, Localazy AI got it right and proposed curita, but the unpredictability of AI shown here is precisely why we still need human review.

A snippet of the brief detailing how to translate the term "Band-Aid".
The original brief indications to translate the term "Band-Aid".
Localazy AI's suggestion for the term "Band-Aid" in Spanish.
Localazy AI's suggestion for the term "Band-Aid" in Spanish.
The key showing the AI-generated context for the term "curita".
The key displayed the AI-generated context for the term "curita".

Sometimes LLMs get things wrong. Most of the time, they are genuinely strong translation tools. What this illustrates is that when the pre-translation analysis layer is in place, linguists have the information they need to make an informed choice, even when the AI suggestion alone wouldn't get them there. The brief and the annotation don't replace AI translation, but they do give both LLMs and human reviewers a better foundation.

It’s easy to get excited by how quickly an AI agent can build something, but basing your whole operational process on vibe-coding can ultimately result in a lot more time and effort than you’d originally think. My suggestion here is just one piece of a much larger puzzle that is still best carried out by dedicated localization software.

💡
AI tools like Claude can be helpful for localization, but basing your whole operational process on them can easily backfire. Check out this article about the downsides of vibe coding your own localization solution and let us know your opinion!

✨ Glossary terms: Surfacing what isn't there 🔗

One of the more unexpected findings came in how Claude handled glossary terms. My test file didn't include a pre-populated glossary but Claude flagged terms in the source text that should be in one. It noted that the reviewer should check whether a given term had a sanctioned translation, or flag it might be worth adding to the glossary if not.

This matters because glossary insertion is one of those features that TMS platforms regularly highlight, but its value depends entirely on the glossary being built over time. What this test suggested is that AI-assisted pre-translation analysis can help identify any glossary gaps before they become mistranslations.

A "verify against glossary" JSON note in a specific string on the Localazy UI.
"Verify against glossary": context notes can also be used for this purpose!

And lastly, in my work with mental health content, clinical review is a must for weighty sections and psychological interventions. Determining which segments qualify for an often pricey local clinical review has largely been a PM task. However, in the brief Claude generated it also flagged sections that could benefit from SME review, which is another huge time-saver.

A "Clinical flag" note for expert review inside a specific string on Localazy.
The notes where expert review was recommended were surfaced in the specific string.

👀 Two tests with very different quality 🔗

With this process in place, I put my tool to the test. I ran the same article through the Localazy’s AI + human review workflow twice: first without context notes, and then with them. What came back were two clearly different outputs:

More natural glossary choices 🔗

In Test 1 (the version without notes), "PTO" was translated literally, rendering the heading as PTO significa Tiempo Libre Remunerado (literally "PTO Means Paid Time Off"), a phrase that keeps an English acronym in a Spanish sentence, leaving readers guessing what the letters even stand for. In Test 2 (the version with context notes in place), the reviewer knew to take some liberties with the PTO line and translated it as Vacaciones = desconectar de verdad ("Vacation means truly disconnecting"), a phrase that captures the spirit of paid time off without forcing an untranslatable acronym into the sentence.

Gender agreement fixes 🔗

Test 1 also slipped between genders when referring to Dr. Zhang, using la Dra. Zhang through most of the piece but switching to el Dr. Zhang in one instance, a small inconsistency but an inconsistency nonetheless. In Test 2, gender agreement for "Dr. Zhang" stayed consistent throughout, most likely because it was flagged.

Attention to loanwords 🔗

Test 1 held onto "mindfulness" as an English loanword, which was fair considering it hadn't been added to the glossary yet. However, Test 2 called this out as a possible glossary term, and the translator committed fully to atención plena, which was the preferred translation and subsequently added to the glossary.

Dialectal adaptations 🔗

The Band-Aid metaphor was also inconsistent in Test 1, appearing once as curita and once as parche (a patch), mangling the metaphor the author was trying to convey. In Test 2, it was translated to parche everywhere it appeared and provided a single thread readers could follow across the piece.

Transcreation space 🔗

Last but not least, Test 1 used skiing as an example hobby, not straying from the English original. On the other hand, after the context notes cleared the section for some transcreation, Test 2 swapped it for fútbol, which would feel far more familiar to a Spanish-speaking audience than a niche winter sport.

After running the two tests, something was clear: the context notes provided a stronger baseline for translation to feel more native

In conclusion, the context notes did more than fix small errors. They provided a stronger baseline and gave the reviewer the right, easily visible tools to prioritize how the content would actually land for the reader. The end result was a version that felt native rather than translated and resulted in a stronger, more engaging global product.

We can see here that AI delivers its real value when it's harnessed within the structure of a dedicated localization tool. Context plus human review consistently outperforms AI working alone, and pairing what AI does best with a solid localization platform is what truly elevates the final result.

🏄‍♂️ Try it out yourself 🔗

You can now use your own pre-generated context notes to match your strings in Localazy. Just upload your JSON file with segment notes (as I did above) and see the magic happen! 🪄 Open any key on the translation interface and they'll pop up to give you and your team an extra context push.

📌 Quick steps for reference here. Contact the Localazy team for any questions and suggestions!

With style guides and glossary configured, review cycles will be reduced to the minimum: AI/MT will pull all the info available at once to suggest translations; humans will get all the context at a glance, without unnecessary back-and-forths; and you'll get quicker, safer content releases.