Security in large language models (LLMs) has become a critical pillar for their enterprise adoption. However, a recent study reveals a concerning vulnerability: alignment mechanisms predominantly trained in English fail when faced with multilingual inputs or code-switching between languages. This epistemic gap allows seemingly robust models to generate harmful responses with full confidence when presented with instructions in low-resource languages. The STEER attack (Safety Targeted Embedding Exploit via Refinement) demonstrates how, through an iterative process of identifying and translating key terms, the model's refusal behavior can be suppressed without losing malicious intent. In tests with 8 billion parameter models, attack success rates of up to 93% were achieved on benchmarks like JailbreakBench, and the generated prompts even transferred to GPT-4o-mini with 35.5% effectiveness. This underscores that the weakness is not architectural but rather one of coverage in training data. For organizations deploying AI for businesses, this finding implies that aligning models in a single language is not enough; cybersecurity strategies are required that include explicit detection of out-of-distribution inputs and continuous multilingual monitoring. At Q2BSTUDIO, we understand that security is not an add-on but a fundamental requirement. That is why we offer custom applications that integrate artificial intelligence with additional verification layers, as well as AWS and Azure cloud services for secure and scalable environments. Furthermore, we combine business intelligence services with process automation and AI agents that reinforce data governance. Our custom software approach allows us to design systems that are not only powerful but also resistant to attacks like STEER, ensuring that the adoption of AI in businesses is as secure as it is innovative.

.jpg)


