我们为最难的长程任务打造自我进化的模型,以及让它们持续变强的专家级引擎:智能体、环境、数据与验证器。
从一个想法到可验证的结果:智能体自主阅读、提出假设、跑实验,把一个真实方法从弱基线一路改进——端到端,全程自动。
面向跨越数小时、需要多次决策的工作流的多步、用工具的智能体,具备长程任务所需的记忆、纠错与技能复用。
每个解决的任务都成为下一步的训练信号。技能沉淀为可复用的技能库;智能体的能力不断累积,而非每次从头开始。
一个自我进化的模型需要完整的循环:驱动它的智能体、它们行动其中的环境,以及训练它的数据与奖励。每一块我们都自己构建,并作为开放研究发布。展开任一类别查看具体工作。
Chat an idea, get a paper, fully autonomous, self-evolving research.
Talk to your agent; it learns and turns conversation into training data.
Two agents co-evolve from zero data via tool-integrated reasoning.
RL agents distill trajectories into a reusable, co-evolving skill library.
Efficient lifelong memory for LLM agents, text and multimodal at ~30x fewer tokens; EvolveMem self-evolves its own retrieval.
Real-time self-evolving VLM agent, frame-gating and skill banks cut API cost dramatically.
Improve a real method from a weak baseline, scored on sealed hidden data.
End-to-end paper reproduction in physics: implement a published method from scratch and match its results, 30 expert tasks across 11 subfields.
Benchmarking AI agents in evolving information environments.
A fully synthetic, database-backed environment generator for agentic RL, train on generated worlds, generalize out of distribution.
Executable interactive benchmarks for command-line agents.
200-scenario benchmark pairing video clips with a persistent workspace and executable checkers.
500 original physics problems, high-school to Olympiad, best model 37% vs humans 62%.
Visual chain-of-thought: models must draw intermediate images to reason.
1.5M+ GPT-4o-refined image-edit triplets for training instruction-based editors.
1.3B web images recaptioned with LLaMA-3 to train CLIP and diffusion models.
25M+ medical images across ten modalities with multigranular annotations.
32,682 medical QA pairs with knowledge-graph reasoning paths for clinical reasoners.
~150K image-question-answer reasoning traces for training R1-style reasoning VLMs.
73K vision-language process-reward samples for training VL reward models.
Safer alignment of reasoning LLMs (e.g. DeepSeek-R1) from just 1K curated examples.
Capability-staged RLVR curriculum, perception, visual-reasoning, and text-reasoning data for VLM post-training.
13.9K medical problems with DAG-structured, knowledge-grounded reasoning traces.
部分发布 · HF 下载量(近 30 天)与 GitHub stars 数据核对于 2026 年 6 月。