Safety in large language models (LLMs) is not uniform across languages. Recent research shows that prompts refused in English can generate harmful responses in low-resource languages. This multilingual jailbreak phenomenon undermines the reliability of AI systems deployed globally. The Minionese benchmark, presented in a mechanistic study, analyzes 18 languages with four perturbation types — standard translation, code-switching, transliteration, and translationese — revealing distinct vulnerability profiles. For instance, transliteration failure depends on script family, while code-switching remains effective even at the lowest resource tiers. A sharp safety regime transition between Tiers 2 and 3 is also observed across all evaluated models.
Mechanistically, low-resource jailbreaks succeed by routing harmful content through a geometric subspace that projects insufficiently onto refusal directions. That is, the refusal mechanism remains intact but is not triggered because the input representation does not align with the safety vector learned during training. This finding implies that English-only safety evaluations are insufficient; enterprises must consider script family, perturbation type, and per-language coverage.
For organizations developing or integrating LLMs, understanding these vulnerabilities is critical. It is not enough to audit a model in English; multilingual applications require exhaustive testing covering minority languages and orthographic variations. This is where custom solutions come in. At Q2BSTUDIO, a software and technology development company, we offer artificial intelligence services that include secure LLM deployment with multilingual assessments. Our team integrates cybersecurity techniques to detect and mitigate these attack vectors, protecting both data and brand reputation.
Furthermore, security architecture must adapt to each client. We work with advanced cybersecurity and pentesting to identify gaps in AI systems, including multilingual jailbreak tests. We also develop custom applications that integrate AI agents with dynamic security controls, and cloud solutions on AWS/Azure that scale on demand. Our Power BI dashboards enable real-time monitoring of content filter effectiveness per language, facilitating data-driven decision-making.
The Minionese benchmark underscores that multilingual safety is not a luxury but a regulatory and ethical necessity. Companies that neglect this aspect risk compliance incidents, data leaks, or reputational damage. By adopting a comprehensive approach — combining automated evaluation, human review, and continuous tuning — robust LLMs can be built. At Q2BSTUDIO, we drive that transformation with AI, cloud, and automation technologies, ensuring every language receives the same level of protection. Mechanistic research gives us a roadmap: understanding how geometric subspaces align or misalign with safety is key to designing systems that truly understand and respect boundaries in any language.





