AI & ML interests

None defined yet.

Recent Activity

omarkamaliย 
posted an update about 1 month ago
view post
Post
183
Apparently grammar can travel without a passport ๐ŸŒ

We trained a new neural grammar model *Tarakib* across 166 languages on top of Sawtone [1] - itself trained on nearly 400 langs. We then tested it on 28 languages that contributed no training updates to the grammar task.

The surprising result: it still learned useful parts-of-speech structure based on the foundational linguistic representation offered by Sawtone.

Performance varied considerably, and the transfer was not limited to close, same-script relatives but crossed writing systems. We found connections across language families, and about a quarter of the stable transfer links crossed writing systems.

This does not mean universal parsing is solved. But it suggests that Sawtone's representations contain universal grammatical clues that a shared model can reuse in languages it was never directly taught to parse.

A small step toward "spaCy for every language" โœŒ๏ธ

1. https://www.omneitylabs.com/models/sawtone
adaamkoย 
posted an update 2 months ago
view post
Post
131
๐Ÿฅฌ LettuceDetect v2 โ€” span-level hallucination detection for code, tool output, and structured documents.

Hallucination detectors are trained on document QA, but agents ground their answers in source code, tool output and markdown. On code-agent answers, existing detectors reach 0.17 span-F1 and even 550B zero-shot judges at most 0.22.

We built a unified span-level benchmark โ€” 74,285 newly constructed examples (145K+ with RAGTruth and 14-language PsiloQA folded in), every span typed and character-labeled โ€” and trained two detectors on it:

๐Ÿค– KRLabsOrg/lettucedect-v2-qwen-2b โ€” generative, typed spans + explanations in one pass, 32K context, **0.689 span-F1** (0.60 on code-agent)
โšก KRLabsOrg/lettucedect-v2-mmbert-base โ€” 307M multilingual encoder for high-throughput setups
๐Ÿท๏ธ KRLabsOrg/lettucedect-v2-taxonomy-head โ€” types the spans of any binary detector

It also reaches the best reported English PsiloQA IoU (0.724) and 81.8 RAGTruth example-F1, so specializing on code didn't cost general RAG performance.

๐Ÿ“š Dataset: KRLabsOrg/lettucedetect-code-hallucination
๐Ÿ“„ Paper: https://arxiv.org/abs/2607.00895

The models are now integrated natively into vLLM Semantic Router โ€” joint blog post on how it works: https://vllm-sr.ai/blog/lettucedetect-v2-generative-hallucination-detection
  • 3 replies
ยท
abidlabsย 
posted an update 4 months ago