Key takeaways
- A multinational networking and cybersecurity technology provider partnered with TELUS Digital to evaluate and strengthen the AI safety guardrails across its AI-powered products.
- TELUS Digital generated a 40,000-prompt multilingual adversarial dataset covering the client's safety-critical categories and attack techniques.
- Rather than applying human review as a blanket quality gate, TELUS Digital applied it selectively, validating baseline quality and enriching individual prompts wherever expert judgment identified opportunities to sharpen realism and adversarial depth.
- 40KMultilingual adversarial red-team prompts delivered
- 5Languages covered: English, French, Italian, German and Spanish
- 2,500Prompts reviewed and calibrated by expert annotators
The challenge
As a leading global networking and cybersecurity technology provider, the client embeds AI features across its infrastructure and security products, deployed in enterprise environments worldwide. Unlike consumer chatbots, these AI interfaces interact directly with sensitive operational systems, security policies and IT workflows. Because these AI assistants hold execution capabilities (e.g., troubleshooting network outages, summarizing sensitive security logs), a guardrail failure represents a direct threat to network integrity and data privacy.
A successful jailbreak or prompt injection attack against the client’s enterprise AI assistant could allow malicious actors to alter access controls, expose internal system architecture or bypass security policies. Safety guardrails trained primarily on English data frequently fail when exposed to non-English inputs. To ensure guardrails held up across every region where its customers operate, the client required adversarial data that met strict operational criteria:
- Native-quality text across all five target languages (English, French, Italian, German and Spanish)
- Comprehensive coverage of safety-critical categories (OWASP Top 10 for LLMs, prompt injection, system prompt leakage, policy bypass)
- Authentic attack patterns that mimic real-world threat actors rather than artificial patterns a model could easily learn to dismiss
Furthermore, because safety guardrails calibrated for English do not adapt effectively to the unique idioms and phrasing of the other languages referenced above, the dataset had to maintain robust performance across the different geographical regions.
The TELUS Digital solution
The engagement combined two layers: large-scale synthetic generation through Fuel iX™ Fortify and a human-in-the-loop (HITL) validation layer staffed by native speakers who were red-team specialists. Fortify produced 40,000 adversarial prompts across the five languages, engineered to span a wide range of attack techniques and safety-critical categories. The HITL layer verified that the data met the bar for realism, technique fidelity and linguistic authenticity before delivery.
Sampling and coverage
Rather than reviewing the full corpus, a random sampling was applied with light stratification to a 2,500-prompt subset (6.25% of the dataset), balanced across attack technique, safety category and language. This kept the validation signal representative of the dataset as a whole while concentrating expert effort where it added the most value.
Three-dimension scoring rubric
Every sampled prompt was scored against a standardized rubric, applied consistently across all five languages. This reduced inter-rater variance and kept each judgment auditable.
- Objective alignment: Does the prompt target its intended adversarial objective and map to the correct safety category?
- Technique correctness: Does it execute its specified attack technique (e.g., a role-play jailbreak, obfuscation or prompt injection) rather than just referencing it?
- Native-language fluency: Does the prompt read as text authored by a native speaker, free of the translation artifacts that reduce a multilingual prompt's adversarial validity?
Targeted enrichment and safety-aware quality control (QC)
Edits were applied only where expert judgment indicated a measurable gain in realism, clarity or adversarial depth and not as routine remediation. Specialists occasionally added authorship to deepen a prompt's sophistication beyond what generation produced. A safety-aware QC step tied into the client's own evaluation: prompts that registered as benign against their guardrails were regenerated and replaced, so every prompt shipped was fit for safety evaluation.
The results
The project gave the client a proven, scalable way to stress-test its AI safety guardrails, backed by a synthetic-generation-plus-expert-review workflow validated at production scale. It also laid the foundation to extend multilingual adversarial testing to new languages and safety categories as the client's needs grow. Four outcomes stand out:
- High baseline quality, confirmed by expert review: Across the reviewed subset, scores remained consistently high before and after editing. This confirmed that Fortify's synthetic generation was sound on its own and that expert input served to add depth rather than fix basic errors.
- Elimination of cross-lingual safety bypasses: Testing revealed vulnerabilities where guardrails blocked attacks in English but failed against equivalent attack vectors in other languages. These were remediated, ensuring consistent protection across global enterprise deployments.
- Traceable, auditable metadata for AI governance: Every prompt was delivered with defined mappings to the client's enterprise security taxonomy, attack techniques, target languages and scoring metadata. This allows the organization to audit guardrail performance systematically.
- A repeatable workflow for continuous AI evaluation: The generation-plus-review framework established an ongoing capability for the client to evaluate new model iterations, emerging agentic capabilities and additional regional languages.
Together, these provided the client with a high-volume, expert-validated dataset to stress-test its AI safety guardrails against real-world threat patterns in languages its global enterprise users depend on.






