Research

Models that improve themselves.

We build self-improving models for the hardest long-horizon tasks, and the expert-grounded engine of agents, environments, data, and verifiers that lets them keep getting better.

What we build

AutoResearch and Self-Improving Agents

AutoResearch

From an idea to a verified result: agents that read, hypothesize, run experiments, and improve a real method from a weak baseline, autonomously, end to end.

General long-horizon agents

Multi-step, tool-using agents for workflows that span hours and many decisions, with the memory, recovery, and skill-reuse that long horizons demand.

A loop that compounds

Every solved task becomes training signal for the next. Skills distill into reusable libraries; the agent's reach grows instead of resetting each run.

How it works · the engine

Agents, environments, and data, all built and open.

A self-improving model needs a full loop: the agents that drive it, the environments they act in, and the data and reward that teach it. We build each piece and ship it as open research. Open a category to see the work.

20k+ GitHub stars collectively 60k+ dataset downloads / month CVPR · ICML · ICLR · NeurIPS
AutoResearch & Self-Improving Agents
Self-improving, omni-modal agents for research and long-horizon work.
Environments & benchmarks
Executable, leakage-sealed environments and benchmarks that score agents against ground truth.
Data, reward & verifiers
Expert-authored post-training data and reward signals, across modalities and domains.

Selected releases · HF downloads (trailing 30 days) and GitHub stars verified June 2026.

It's all open.

Our agents, environments, and datasets are public. Browse the repos, or talk to us about your model.

Request access