Recursive self-improvement
What forms of self-refinement, self-evaluation, and autonomous research loops produce real improvement rather than cosmetic optimization?
Recursive self-improvement, measured against reality.
DeepGrounding is an independent research organization studying how self-improving AI systems can remain honest, calibrated, and useful as their planning horizons lengthen.
Research program
What forms of self-refinement, self-evaluation, and autonomous research loops produce real improvement rather than cosmetic optimization?
How should we evaluate systems that operate over many steps, revise their own plans, and accumulate hidden failure modes over time?
Which external signals, tests, audits, and verification protocols keep self-improving systems connected to truth instead of their own preferences?
Publications
Format repair can masquerade as self-correction at small-to-mid scale: reported gains from LLM self-revision are often driven by an answer becoming extractable, not by better reasoning. We decompose the effect, test it causally, and show format effects dominate at capable model scale while genuine content-level change concentrates — and is often harmful — at floor scale.
A survey of 1,547 arXiv papers (2024–2026) mapping the field's response to the gap between single-step model capability and reliable long-horizon task completion, across planning, memory, execution, training, and evaluation.
Operating principles
DeepGrounding treats recursive self-improvement as an empirical measurement problem: a system should not receive credit for improvement unless its gains survive independent checks.
Current work focuses on calibration floors, evaluator drift, verifier circularity, and the gap between self-reported progress and externally grounded performance.
Publications and their reproducibility bundles are released as each manuscript clears review. See Publications for what's out so far; more research pages are in preparation.
Contact
Email deepgroundingai@gmail.com or follow the work on GitHub.