We build self-improving models for the hardest long-horizon tasks, and the expert-grounded engine of agents, environments, data, and verifiers that lets them keep getting better.
From an idea to a verified result: agents that read, hypothesize, run experiments, and improve a real method from a weak baseline, autonomously, end to end.
Multi-step, tool-using agents for workflows that span hours and many decisions, with the memory, recovery, and skill-reuse that long horizons demand.
Every solved task becomes training signal for the next. Skills distill into reusable libraries; the agent's reach grows instead of resetting each run.
A self-improving model needs a full loop: the agents that drive it, the environments they act in, and the data and reward that teach it. We build each piece and ship it as open research. Open a category to see the work.
Chat an idea, get a paper, fully autonomous, self-evolving research.
Talk to your agent; it learns and turns conversation into training data.
Two agents co-evolve from zero data via tool-integrated reasoning.
RL agents distill trajectories into a reusable, co-evolving skill library.
Efficient lifelong memory for LLM agents, text and multimodal at ~30x fewer tokens; EvolveMem self-evolves its own retrieval.
Real-time self-evolving VLM agent, frame-gating and skill banks cut API cost dramatically.
Improve a real method from a weak baseline, scored on sealed hidden data.
End-to-end paper reproduction in physics: implement a published method from scratch and match its results, 30 expert tasks across 11 subfields.
Benchmarking AI agents in evolving information environments.
A fully synthetic, database-backed environment generator for agentic RL, train on generated worlds, generalize out of distribution.
Executable interactive benchmarks for command-line agents.
200-scenario benchmark pairing video clips with a persistent workspace and executable checkers.
500 original physics problems, high-school to Olympiad, best model 37% vs humans 62%.
Visual chain-of-thought: models must draw intermediate images to reason.
1.5M+ GPT-4o-refined image-edit triplets for training instruction-based editors.
1.3B web images recaptioned with LLaMA-3 to train CLIP and diffusion models.
25M+ medical images across ten modalities with multigranular annotations.
32,682 medical QA pairs with knowledge-graph reasoning paths for clinical reasoners.
~150K image-question-answer reasoning traces for training R1-style reasoning VLMs.
73K vision-language process-reward samples for training VL reward models.
Safer alignment of reasoning LLMs (e.g. DeepSeek-R1) from just 1K curated examples.
Capability-staged RLVR curriculum, perception, visual-reasoning, and text-reasoning data for VLM post-training.
13.9K medical problems with DAG-structured, knowledge-grounded reasoning traces.
Selected releases · HF downloads (trailing 30 days) and GitHub stars verified June 2026.
Our agents, environments, and datasets are public. Browse the repos, or talk to us about your model.