The new accountability problem
AI can now draft, translate, classify, summarize, and publish content faster than any human review team can keep up with. In most enterprises, Large Language Models (LLMs) and Neural Machine Translation (NMT) are no longer isolated experiments. They are wired directly into Content Management Systems, translation workflows, customer-support tools, and release pipelines.
The productivity case is real: Content moves from source to market in minutes instead of days, and global releases no longer wait for every language to complete the same sequence of hand-offs. But automation also removes something that used to be built into the process by default: friction. A technical writer, a Subject Matter Expert, a translator, and a language editor each represented a chance to catch a factual, contextual, cultural, or terminology error before it reached a customer. When those checkpoints disappear, the organization has to replace them deliberately – or accept that no one is replacing them at all.
Without governance | Review everything, every time – or accept blind deployment at scale. |
With governance | Route content according to risk, and review exactly what needs human judgment. |
The question enterprises now face is no longer simply “Can we automate this?” It is “Under what conditions is automation safe, and who owns the decision?”
AI does not create enterprise content risk – ungoverned decisions and ownerless pipelines do.
Why the most dangerous failures look correct
The obvious AI error is easy to spot: broken syntax, missing text, a wrong term, a visibly poor translation. The more dangerous failure is the one that looks correct. Such cases are referred to as “Quiet Failures.”
Generative AI models are optimized to produce plausible output – they are probabilistic prediction engines, not deterministic fact engines, and understand syntax probability better than truth. This is what creates “fluent nonsense”: grammatically flawless, highly polished text that is confidently wrong. It triggers a “fluency bias” in human reviewers – the better an AI draft reads, the less likely a reviewer is to doubt the underlying facts or source fidelity. The EU AI Act explicitly recognizes this kind of automation bias as a risk that human oversight measures should account for in high-risk AI systems. [1]
A hypothetical “Quiet Failure”
Consider a long-lifecycle product in energy, aerospace, or medical technology. An automated pipeline is generating localized safety documentation from a repository containing years of legacy material. The output uses the correct structural tags, approved terminology, and a spot-on tone.
But the model has pulled a critical safety-clearance measurement from a 2018 specification instead of the current 2026 source text. Nothing in the sentence signals the error – to the model, the older number was simply more “statistically probable,” because the 2018 spec appeared in roughly 500 legacy documents in the training data, while the 2026 spec appeared in only five. To an overworked engineer skimming a 200-page document, the text passes spellcheck and grammar checks without a flag.
This is a Quiet Failure: The defect is not linguistic fluency but source fidelity. Traditional QA can miss it because the sentence is grammatically correct and internally coherent. Catching this class of error requires controls that test provenance, version, terminology, and business rules – not just surface quality.
Shadow AI: The second governance gap
The production pipeline is only half the picture. When sanctioned tools are slow, hard to access, or too restrictive for urgent work, employees under pressure turn to public AI services on their own – pasting unreleased specifications, source code, or internal HR reviews into a free LLM for a “quick summary” or “fast translation.” This “Shadow AI” can move sensitive material outside the organization’s controlled environment before security or compliance teams ever know it happened.
According to TELUS Digital’s 2025 AI at Work survey, nearly 68% of enterprise employees who use GenAI at work reported accessing public assistants through personal accounts rather than company-approved platforms, and 57% said they had entered sensitive information into those tools. [2] The figures come from a single vendor survey and should be read as an indicator of behavior, not a universal estimate of the workforce. Yet, the direction is consistent with what most enterprise security teams already suspect: When proprietary data enters a public-tier LLM, it can be absorbed into model training sets, risking permanent data exposure, regulatory non-compliance under GDPR, and IP loss.
Therefore, governance has to address both sides of AI use: how systems generate and publish content and how people are permitted to interact with AI tools. Policy matters, but so do usable enterprise alternatives, secure gateways, data controls, employee training, and a clear escalation path when employees need help getting something done quickly.
Four warning cases – and what they actually reveal
Several public cases illustrate the accountability problem, but they deserve to be read carefully.
Case 1: Starbucks Korea, 2026
Starbucks Korea launched a promotion for a "Tank" tumbler line on May 18 (the anniversary of the 1980 Gwangju Uprising) with a slogan that echoed language tied to a notorious 1987 custody death. [3] An internal investigation found that the campaign concept and slogan had been generated with an AI tool, and that the historical date had "never crossed the minds" of a marketing team that judged the launch day based on commercial data alone. The campaign read as polished and on-brand; nothing in the copy itself signaled a problem. It passed through several rounds of approval before reaching the public. Shinsegae Group fired the Starbucks Korea CEO, and its chairman issued a public apology. This is a Quiet Failure in marketing copy rather than technical content: fluent, catchy language that no one in the approval chain stopped to question for context.
Case 2: Air Canada, 2024
In Moffatt v. Air Canada, the BC Civil Resolution Tribunal found Air Canada responsible for incorrect information supplied by the chatbot on its website, rejecting the airline’s argument that the bot was a “separate legal entity.” The tribunal ruled that Air Canada “is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot.” [4] A company cannot treat a customer-facing chatbot as detached from its own published service.
Case 3: Mata v. Avianca, 2023
A U.S. federal court sanctioned lawyers after they submitted non-existent case citations generated by ChatGPT and failed to verify them. [5] The lesson is not that AI is prohibited – it is that professional gatekeeping responsibilities stay with the human organization using the tool.
Case 4: Amazon Sweden, 2020
During Amazon’s Swedish launch, automated translation produced confusing and offensive product listings. [6] The incident became a widely cited example of how machine translation without contextual and linguistic oversight can create reputational problems at scale.
None of these cases creates a universal rule that a company is automatically liable for every AI output. What they show is more practical: Organizations cannot assume that responsibility disappears just because a machine was involved.
Make accountability explicit
Governance becomes practical when ownership is visible. Technical communication and localization teams are well positioned to act as Risk Architects: the people who help decide where automation is sufficient, where human judgment is mandatory, and what evidence is required before content is released. This does not mean moving all accountability into the content function – it means distributing decisions across the teams that actually control context, rules, infrastructure, and release. No critical workflow step should be “owned by the pipeline” simply because the pipeline is automated.
Stakeholder roles
- Technical communication & localization (Risk Architects): Design automated quality gates, manage context repositories (glossaries, CCMS metadata), and establish circuit breakers
- Product & marketing (Context Owners): Define content intent and business impact before feeding material into automated engines
- Legal & compliance (Boundary Setters): Codify non-negotiable disclaimers and safety boundaries directly into validation rules
- IT & AI operations (Infrastructure Custodians): Secure closed-loop enterprise gateways to prevent data leaks and maintain model stability
The RACI matrix
A structured RACI model – defining who is responsible, accountable, consulted, and informed – makes those boundaries explicit. The exact allocation will differ for each organization; the important point is that every workflow step has a named owner.
Workflow stage | Product / | Tech comm & | Legal & | IT & AI |
Risk tier assignment | Accountable | Responsible | Consulted | Informed |
System prompts & glossaries | Consulted | Accountable | Informed | Responsible |
Secure AI infrastructure | Informed | Consulted | Informed | Accountable |
Compliance & legal rules | Informed | Responsible | Accountable | Informed |
Quality gate thresholds | Consulted | Accountable | Consulted | Responsible |
Human-in-the-loop execution | Informed | Accountable | Informed | Informed |
Incident response & audit | Responsible | Responsible | Accountable | Responsible |
The shift also changes what content teams should measure. “Words per hour” or total production volume says little about governance effectiveness. More useful measures include how often automation is interrupted, how many high-impact errors are caught before publication, and where human review adds the most value.
The shift in value | |
Old metric | Words per hour/production volume |
New metric | Risk mitigated/catch rate (e.g., errors caught per 100 pages) |
Route content by risk, not by habit
A common response to AI risk is to review everything. But this works only until volume makes full review impossible. A better strategy combines the potential impact of an error with the likelihood that the workflow will actually produce the error. The three-tier model shown in Figure 1 is an operational starting point, not a legal classification. To make it functional, it needs to be adapted to your products, markets, and regulatory obligations.
The key principle is this:Low linguistic complexity does not automatically mean low business risk. Risk classification should weigh consequence, audience, reversibility, regulatory exposure, and the quality of the source context, not just how simple the sentence looks.
Put the controls inside the workflow
Risk tiers only matter when they lead to concrete control points. For Tier 2 (medium-risk) content, an automated “circuit breaker” can stop a publication when defined conditions are not met: Generate a draft, run deterministic checks for terminology and structure, run a quality-estimation or secondary model check, then either publish automatically or route the item to human review.
Tools such as COMET or an LLM-based evaluator are useful components, but neither should be treated as a universal truth detector. They estimate or assess aspects of output quality; they do not establish factual correctness on their own. Thresholds should be calibrated against your own content and risk tolerance, not borrowed from a vendor default.
A practical roadmap for enterprise governance
An enterprise AI content governance program can be set up in the following four phases:
- Map the current state: Inventory AI and NMT pipelines, identify where content enters and exits automated systems, and look for Shadow AI practices.
- Define ownership and risk: Create a cross-functional governance group and assign RACI responsibilities. Classify content by impact and likelihood, not simply by format or department.
- Add control points: Implement secure gateways, deterministic checks, quality estimation, approval rules, logging, and explicit circuit breakers for high-impact failure modes.
- Measure what the controls prevent: Track human intervention, escaped defects, critical errors caught before publication, and trends in the types of corrections humans make.
Change the conversation
Executive reporting often focuses on volume and speed. Governance adds a second question: What risk did the automation prevent or expose? Two measures are particularly useful: The Human Intervention Rate (HIR) shows how much automated output still requires correction. The Catch Rate shows how many important errors or compliance violations are caught before publication. Tracked over time, these measures reveal whether automation is getting more reliable, whether human review is focused on the right content, and where controls still need work.
A Monday morning red-flag checklist
Use this checklist to spot high-risk “Quiet Failures” before they reach a customer.
Perfect-looking reference: Does the output cite a part number, specification, policy, or source that cannot be verified against the current repository?
Ghost specification: Is a term, measurement, or requirement coming from retired or obsolete documentation?
Vague safety language: Has a precise requirement been replaced with a generic phrase such as “ensure proper safety” instead of an explicit value like “torque to 15 Nm”?
Prohibited claim: Has the system introduced an unsubstantiated adjective – for example, “fastest” or “safest” – that the organization cannot support?
Stripped metadata: Has AI removed DITA attributes, hazard tags, identifiers, or structural information needed downstream?
Missing ownership: Can the team name the person or function that can stop publication when the automated control fails?
Governance is becoming part of the operating environment
The governance model described here is not a substitute for legal compliance, and its three content tiers are not legal categories. But the direction of regulation reinforces the value of explicit human oversight. Article 14 of the EU AI Act requires effective human oversight for high-risk AI systems. It states that oversight measures should be proportionate to the risks, level of autonomy, and context of use. It explicitly calls out automation bias and the need for people to be able to override or stop the system where appropriate. [1]
For technical communicators and localization professionals, the implication is practical: documenting who checks what, under which conditions, and with what authority is increasingly part of responsible AI operations, not an optional extra.
Governance as the accelerator
The goal of enterprise AI governance is not to return to manual review of every word. It is to make selective human judgment more effective. That means treating content as a governed production system: Define the source of truth, classify the risk, assign ownership, place controls at the points where failures can be caught, and preserve an audit trail for decisions that matter.
The strategic value of the content function is changing as a result. It is no longer measured only by how much content can be produced or translated; it is also measured by how confidently an organization can automate without losing control of what it publishes.
Technical communication and localization teams are well positioned to lead that shift. By becoming Risk Architects rather than simply final-stage reviewers, they can turn AI from an unpredictable drafting engine into a scalable – and defensible – content capability.
References
[1] European Union, Regulation (EU) 2024/1689 (EU AI Act), Article 14, Human oversight; consolidated text accessed September 2026.
[2] TELUS Digital (2025), “TELUS Digital Survey Reveals Enterprise Employees Are Entering Sensitive Data Into AI Assistants More Than You Think,” 26 February 2025.
[3] Reuters and PRovoke Media reporting on Starbucks Korea's "Tank Day" campaign, May–June 2026; Yonhap News Agency, internal-investigation findings on the campaign's AI-assisted origin.
[4] Moffatt v. Air Canada, 2024 BCCRT 149; see also American Bar Association, “BC Tribunal Confirms Companies Remain Liable for Information Provided by AI Chatbot.”
[5] Mata v. Avianca, Inc., No. 22-cv-1461 (PKC), U.S. District Court for the Southern District of New York, 22 June 2023.
[6] The Guardian, “Amazon hits trouble with Sweden launch over lewd translation,” 29 October 2020.



