Khalid Yusuf Dahir
Founder, Unkad Labs
Safety properties in language models are learned, and therefore fragile. I measure where they break: across languages, across evaluators, and across the gap between what a model does and what its overseers can see.
- Running Qor Af-Soomaali, a community corpus platform with 159 contributors and personal weekly goals that ladder to a 100,000-sentence campaign.
- Publishing research notes at unkad.com: six so far, each with open code, raw judgements, and BibTeX.
- Preparing SomaliBench for peer review: separating genuine harmful compliance from low-quality generation.
Multilingual safety · Unkad Labs
Do the safety properties of AI survive a change of language?
Llama 3.1 refuses 97% of harmful requests in English and 7% in Somali; open safety filters catch 100% in English and as little as 6% in Somali.
Alignment and oversight
What do overseers miss when models adapt to being watched, or content falls outside what a judge can read?
A model trained against a blind-spotted lie detector kept its deception machinery and stopped using it only where the watcher could see; an LLM judge denied a readable source collapses from a 62% to a 6% yes rate rather than becoming uncertain.
The lie moved. The liar did not.·The overseer that cannot read