Submitted by Qihao Zhao 63 ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes Microsoft 2.67k 3
Submitted by yangyu huang 64 ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog Microsoft 7
Submitted by Amirhossein Abaskohi 11 SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference Microsoft 4 2
Submitted by qianchu liu 8 HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents Microsoft 47 2
Submitted by Yanuo Ma 9 Building to the Test: Coding Agents Deliver What You Check, Not What You Requested Microsoft 0 2
Submitted by Shaoqiu Zhang 96 FastContext: Training Efficient Repository Explorer for Coding Agents Microsoft 6
Submitted by Tejas Agrawal 1 A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets Microsoft 1 3
Submitted by Wanli Li 108 WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Microsoft 160 2
Submitted by Rui Yang 23 OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents Microsoft 46 3
6 Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory Microsoft 12