EverMind Open-Sources Raven V0.2.0 Multi-Agent AI Framework for Harness-Layer Self-Improvement
An experimental Curator module lets AI rewrite its own operating behavior without touching model weights.

EverMind, an AI company incubated by Shanda Group, released Raven V0.2.0 on September 23, 2026 — a significant revision of its open-source multi-agent AI framework that the company describes as the "Harness of Harnesses," a system designed to orchestrate specialized AI agents, preserve cross-session experience, and improve its own operating procedures through feedback, without retraining the underlying models. The release carries an Apache-2.0 license and is available on GitHub.
The framework's central argument challenges how the field normally thinks about recursive self-improvement. Most discussions of AI self-improvement focus on model weights — the internal parameters that encode what a language model knows and how it reasons. EverMind draws instead on the Complementary Learning Systems theory from cognitive neuroscience: the hippocampus quickly absorbs specific experiences while the cortex slowly consolidates general knowledge. In Raven's framing, model weights are analogous to cortex — good for slow consolidation — while the harness layer, which governs how a model plans, uses tools, manages memory, and executes actions, should work like the hippocampus: fast to adapt and accessible to AI-driven modification.
The upshot is a framework that separates what an AI model knows from how it works. Raven V0.2.0 is designed so that AI can rewrite the second category, updating prompts, tool availability, strategic code, and playbook structure without touching the first.
Read more: RSI: AI Accelerating AI — the bottleneck has shifted from capability to verification
The Harness of Harnesses: Orchestrating Specialist Agents
Raven's architecture is built around a two-layer structure. At the outer layer, a host agent — Raven itself — accepts a complex goal, decomposes it into a directed acyclic graph of tasks, assigns each node to a specialist agent, and manages dependencies, handoffs, and results. At the inner layer, each specialist agent brings its own execution mechanism, tooling, and memory to its assigned work.
Raven V0.2.0 ships with four built-in specialist agents: Raven-Research, for deep research and literature analysis; Raven-Code, for coding, debugging, and data analysis; Raven-Design, for presentation creation, web pages, visual deliverables, and brand materials; and Raven-Oncall, for long-running experiments and continuous execution tasks. A fifth agent, Raven-PPT, ships as a hidden engine that Raven-Design routes PowerPoint work to.
External agents — including Claude Code, OpenAI's Codex, and others — connect through the Agent Connector, which supports the Agent Communication Protocol, standard command-line interfaces, and OpenAI-compatible APIs. The connector tests each agent with a live prompt before marking it available, so a connection failure or authentication problem surfaces immediately rather than mid-task. Once connected, an external agent receives task assignments and returns deliverables the same way a built-in agent does; the orchestration layer treats them equivalently.
The practical consequence is that a single Raven task graph can route a research subtask to Raven-Research, a coding subtask to Claude Code, and a design subtask to Raven-Design, with the output of each feeding into the next. EverMind demonstrated this in a cantilever beam simulation scenario: Raven-Research identified methods, Raven-Code implemented the simulation, and Raven-Oncall ran the experiment and monitored convergence results. The task graph tracked which upstream outputs each downstream node required, skipping dependent nodes when a predecessor failed and continuing independent branches.
Playbooks — saved task graphs the user can store and reuse — let teams capture a working multi-agent workflow and invoke it again by name. EverMind frames this as an alternative to writing automation scripts: the playbook is built interactively the first time a task runs, then made available for subsequent invocations without requiring a developer to formalize it as code.
The Curator: AI Rewrites Its Own Operating Layer
The more architecturally novel addition in V0.2.0 is the Curator, shipped as an experimental feature in the repository under experimental/curator. The Curator is a component that reads the current harness configuration, reviews task outcomes and user feedback, generates proposed changes to four categories of harness elements — prompt and playbook content, tool availability and gate policies, strategy code, and module composition — and installs those changes after a validation cycle.
EverMind structures the modifiable elements of the harness across four strategy interfaces: Memory (what context a given conversation can access), Planning (the sequencing logic that determines task order), Capability (which tools the model can call at any point), and Action (what checks happen before an action executes). The Curator generates Python class implementations of these interfaces, binds them to existing callback points, and runs pre-install checks before writing anything to disk. If installation fails, the previous harness version restores automatically.
No source-code modification of Raven's core is required. All changes land in extension points EverMind designed for AI-driven modification: the agent's home directory, plugin slots, and public strategy protocol interfaces.
To demonstrate the Curator's operation, EverMind published a simulated case study — experimental/simulation/cases/s0925c — in which a fictional travel agency owner provided materials, played the role of a customer, and gave feedback after each simulated session. The Curator generated process code and rubric checks from the initial materials alone, and then converted user feedback into specific adjustments to the executing harness over multiple rounds. By the third evaluation round, all 11 red-line criteria in that evaluation passed. EverMind is explicit that this is one simulated case, not a statistical validation, and that the Curator remains experimental. Cross-scenario robustness testing and integration with model-weight-level updates have not yet been completed.
Read more: OpenAI chief scientist warns AI monitoring is eroding before RSI arrives
Benchmark Results: Vendor-Reported Performance Against Named Comparisons
EverMind published a technical report alongside the release, dated September 27, 2026, covering benchmarks across orchestration, research, coding, design, and long-running experiment tasks. All figures below are from EverMind's own evaluations unless otherwise noted; independent replication has not been confirmed.
On multi-agent orchestration — measuring how accurately the system generates task graphs — Raven reported the following against Hermes and Claude Code under the same base model. Using Qwen3.8-27B as the base, Raven scored 0.923 on Node F1, 0.812 on Edge F1, 0.950 on Partial Order Accuracy, and 0.711 on Exact Match Rate; with DS-V4-Flash-0731 as the base, the corresponding scores were 0.963, 0.897, 0.918, and 0.867. EverMind reports both as above the compared systems under the same model.
On deep research benchmarks combining BrowseComp, FRAMES, HLE, and xBench-DeepSearch, Raven-Research scored 56.3% with Qwen3.6-35B, 59.3% with Qwen3.5-397B, and 76.5% with DS-V4-Flash, compared to a DeepSeek-Harness baseline of 49.7%, 56.0%, and 68.9% respectively.
On SWE-bench Pro — a coding benchmark for resolving real software issues — Raven-Code using Qwen3.8-27B reported a 54.4% problem-resolution rate versus 52.4% for Claude Code. On SWE-bench Verified with DS-V4-Flash, Raven-Code reported 91.0% versus 90.4% for DeepSeek-Harness and 90.2% for OpenCode. On SWE-Refactor, a code-quality benchmark, Raven-Code scored 16.5 versus Claude Code's 7.0 in the DS-V4-Flash configuration; with GPT5.6-Luna Max, Raven-Code scored 13.5 versus Codex's 10.5. On the DataAgentBench live leaderboard, Raven-Code using Opus-5 reported a Pass@1 of 0.8762 versus 0.8713 for the next-ranked system.
On a long-running experiment task — an AI4AI evaluation using a Nanochat 50M pre-training task on two A800 80GB GPUs — Raven-Oncall achieved a 5.8% relative reduction in validation bits-per-byte (val_bpb, a measure of language model perplexity where lower is better) across seven rounds and 172 training runs, within a fixed per-run compute budget of 20 minutes per GPU.
On an internal AI4S scientific benchmark, Raven-Oncall using Opus-5 reported an 82.35% task success rate versus 64.71% for Claude Code with the same model.
RSI in Practice: AI Improving AI's Research Methods
Separately from the Curator, EverMind ran a Raven RSI experiment in which Raven autonomously conducted pre-training experiments, using the val_bpb metric as the objective it was not allowed to modify. Raven planned each of seven rounds, wrote training code, launched experiments, and chose adjustments based on results. The 5.8% relative improvement over 172 training runs, with no training crashes, is the headline outcome.
EverMind also used Raven to produce the promotional materials for the V0.2.0 release itself: a browser-based physics game, a 16-page product deck, Chinese and English posters, and a project README. The company presents this as a demonstration that the four specialist agents can combine to cover a complete production workflow — code, content, visual, and documentation — within a single task graph.
Platform Details
Raven V0.2.0 introduced a web UI served by raven web, replacing the previous terminal-only interface. The UI includes a conversation view, a settings dialog covering providers, skills, plugins, channels, memory, and scheduling, an Agent Connector for managing external agent connections, and a desk with file, agent, and task panes. A Docker image and Compose file ship with the repository, enabling container-based deployment. The installer now supports Linux, macOS, WSL2, and native Windows.
EverOS, EverMind's memory layer, provides cross-session persistence for user context, agent experience, and world knowledge. SkillForge retrieves relevant skills from a local library, EverOS memory, and a SkillHub catalog on demand. Both are included in the default installation.
EverMind plans to launch a cloud-hosted version of Raven through its EverMe product in October 2026.
The full codebase, installation instructions, documentation, and the Curator simulation case are available at github.com/EverMind-AI/Raven.