Tous 674 🔐 Cybersécurité 415 🤖 Intelligence Artificielle 253 💻 Tech & Transformation Digitale 6
Chargement…
🧠
Intelligence Artificielle

Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and…

Lire l'article
🧠
Intelligence Artificielle

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

arXiv:2608.14626v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper, we conduct a Systematic Literature Review (SLR) of LLM safety alignment in low-resource languages…

Lire l'article