Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Priyanka Balodi
KnowDis
2
1
Follow
0 followers
Β·
5 following
AI & ML interests
None yet
Recent Activity
liked
a model
23 days ago
nvidia/Nemotron-Labs-Audex-30B-A3B
reacted
to
albertvillanova
's
post
with π₯
23 days ago
π KTO is now part of the stable TRL API As of Promote KTO to stable API, KTOTrainer and KTOConfig have graduated from trl.experimental to the stable trl API. https://github.com/huggingface/trl/pull/6175 This one closes out a long road. Over the past 6+ months, the "Align KTO with DPO" effort landed ~90 PRs methodically bringing KTO up to the standard we hold for stable trainers, one carefully-scoped change at a time: - Feature parity with DPO: full VLM support (incl. multi-image), sync_ref_model, PEFT + Liger, ZeRO-3 + PEFT dtype fix, pad_to_multiple_of, activation offloading, IterableDataset and dict eval_dataset, remove_unused_columns, and reference-logprob precomputation at init. - Consistency with DPO: aligned method order and signatures, tokenization, _prepare_dataset, PEFT handling, ref-model preparation for distributed training, and config layout β plus a new DataCollatorForKTO and output format. Metrics moved into _compute_loss and simplified to direct averages via the shared _metrics attribute. - Removing legacy baggage: dropped encoder-decoder support, BOS/EOS handling, null_ref_context, generate_during_eval, model_init, preprocess_logits_for_metrics, model/ref adapter names, and several dead config knobs. - Coverage: a full test suite mirroring DPO, text collator tests, VLM tests, and slow tests. - The promotion itself: the experimental β stable move (#6175) and shim cleanup (#6287), handled so downstream users get a clean deprecation path. Honestly, this has been one of the more complex tasks I've taken on since joining the team, not because any single change was hard, but because it demanded sustained consistency across a ~2,000-line trainer, with every branch, comment, and edge case kept in lockstep with DPO. Huge thanks to everyone who reviewed along the way (especially @qgallouedec), the incremental review cadence is exactly what kept this maintainable. KTO now sits on equal footing with our other flagship trainers. π
upvoted
an
article
4 months ago
Multimodal Embedding & Reranker Models with Sentence Transformers
View all activity
Organizations
None yet
spaces
1
Build error
Agents
Trackio
π
models
0
None public yet
datasets
0
None public yet