An assessment, not a release note

Governed Habitats: What the Agent-Habitat Literature Asks of a Harness, and Where HANRIA Stands

HANRIA is a coined name. No expansion, derivation, or translation is published or intended. A U.S. trademark application is pending; no registration or completed clearance is claimed.

I. From Habitat to Institution

Savva et al. (2019) presented Habitat as a platform for training embodied agents. Its project page defines a habitat as the place "where AI agents live and learn" (Meta AI n.d.). This article uses the term in that sense, extended from embodied simulators to any controlled environment in which agents are trained, evaluated, and run.

Inference: that literature spends its detail on measurement: speed, safety, reproducibility, and benchmark scores. Governance arrives in a separate body of work, largely as practices an operator may choose to adopt.

The governance problem is older than the technical usage. Lucasfilm's Habitat, an online virtual world launched in 1986, ran into governance problems with its users. Players exploited bugs and modified their clients. The designers concluded that "detailed central planning is impossible," and when users wanted order, the designers built a voting mechanism and held an election for a Sheriff (Morningstar and Farmer 1990). Inference: the lesson was institutional, not technical. A habitat that hosts autonomous actors needs rules.

This article reads the HANRIA harness against that lesson. HANRIA's build notes put a rule layer before every action (HANRIA build notes 2026). Inference: that ordering invites a different reading of the habitat. It becomes the set of actions the rules admit, not a place agents are dropped into. This article proposes that reading as a classification, not an analogy: a governed habitat is an institution. It judges the harness by three standards this article proposes as its own test: enforcement on the only path to action, a record that an insider cannot rebuild without leaving a trace, and validation by parties who do not control the system. No single source states them as a set.

Status. The HANRIA harness is in development and has not been deployed. Changes go through internal testing and review only. The harness has not been independently audited (HANRIA build notes 2026). The public site says the same of the wider system (HANRIA 2026a; 2026b; 2026f). Every claim below is read against that status.

Evidence. Claims about the literature cite the primary papers. Claims about HANRIA cite either internal build notes from September 2026 or public pages on hanria.ai. The public pages do not mention the harness, so every harness-specific fact rests on the build notes alone and is single-source. Conclusions that go beyond the cited sources are marked Inference. The thesis itself, the reading of a governed habitat as an institution, is this article's argument throughout.

The argument runs through five themes: simulated and sandboxed environments, agent benchmarks, containment, governance and auditability, and external validation. Each theme is scored the same way: where HANRIA aligns with the literature, where it goes further, and where it falls short. Each gap is paired with a proposed fix.

II. Scope

"Habitat" carries three readings in the sources for this article. The first is the environment in which agents operate. The second is ecological, as in the literature on species distribution models (Elith and Leathwick 2009). The third is Lucasfilm's Habitat. This article concerns the first and uses the third as governance precedent.

The simulator literature gives the following grounds for controlled habitats. Training in the physical world is slow, dangerous, costly, and hard to reproduce. Simulation is fast and safe, and it makes benchmarking fair (Savva et al. 2019). Inference: this article treats speed, safety, and reproducibility as the baseline any governed habitat must also meet.

III. The Object of Analysis

HANRIA is a Rust platform that runs AI agents on the owner's own machines under rules. According to the build notes, a rule layer must admit every action before it runs. Each action is recorded in a tamper-evident, hash-chained log, and each agent has its own view of what it did and what that cost (HANRIA build notes 2026). Whether the unreleased enforcement runtime is what performs that admission is the question Section VI takes up.

The platform has three layers. The substrate decides what may run and keeps the record. The cockpit shows what did run and what it cost. The harness runs models under governance on top of both. It depends on pinned versions of the substrate and cockpit and never edits either one (HANRIA build notes 2026).

The harness is staged. Version 0 wraps a fixed model with authorized-only tools, checkpoint and resume workflows, and encrypted traces indexed through the cockpit. It uses only synthetic or public practice tasks. Models are drawn from a local panel of about 40 Ollama models on two NVIDIA GB10 nodes, plus Mac models (HANRIA build notes 2026). The notes describe a version 1 in which independent reviewers would check the work, each putting reputation on the line (HANRIA build notes 2026). The how-it-works page describes reputation-backed review for the wider system and does not mention the harness (HANRIA 2026b). The notes describe a version 2 that would add a sandboxed self-optimizer whose changes require an audit PASS and the owner's acceptance (HANRIA build notes 2026). The harness has not been deployed, and the sources do not say that version 1 or version 2 has been built.

IV. Simulated and Sandboxed Environments

Aligns. Version 0 uses only synthetic or public practice tasks (HANRIA build notes 2026). Savva et al. (2019) justify simulation because training in the real world is dangerous, costly, and hard to reproduce. Inference: version 0's practice tasks apply the same caution to software agents, and they limit exposure while behaviour on the owner's real data is unknown. That application is this article's. Savva et al. write about embodied training.

Gap: evaluation realism. Benchmarks set on self-hosted websites and on real operating systems report wide human-agent gaps. On WebArena, which self-hosts websites in four areas, the best GPT-4 agent reached 14.41 percent end to end, against 78.24 percent for humans (Zhou et al. 2024). On OSWorld, which runs real Ubuntu, Windows, and macOS machines across 369 tasks, the best model reached 12.24 percent, against 72.36 percent for humans (Xie et al. 2024).

Inference: these gaps are consistent with realism making tasks harder, though the sources do not isolate realism as the cause. Results on practice tasks may not predict how an agent behaves on the owner's machines with real data. A safe habitat is not the same as a representative one.

Fix. Inference: build a graded ladder of synthetic tasks, each more realistic than the last, and require agents to clear it before deployment. The ladder keeps practice data synthetic while showing where behaviour changes as tasks grow harder. It would not close the gaps WebArena and OSWorld report, and clearing it would not by itself show how an agent behaves on the owner's real data.

V. Agent Benchmarks and What the Tests Measure

What exists. Harness commits pass an exact-commit gate. The Fedora 44 and AlmaLinux 10 runs passed every test, 307 to 308 tests per run. The payload manifest was ported to Rust with 13 golden tests. Other commits added mutation-testing checks, a confidentiality-first fix, and fixes for stale sockets and pipe draining. The build notes name Fedora 44 on an HP ZBook as the production host and AlmaLinux 10 as a required compatibility check (HANRIA build notes 2026). "Production" names the target host. It does not mean the harness has been deployed.

Classification. Inference: these are software-assurance tests of the harness itself, not measures of how well agents perform inside it. They show that the checked tests passed, not that the harness does everything its code says. The two kinds of evidence must be reported separately, or test passes will be read as capability.

Not yet aligned. Kapoor et al. (2024) find that agent benchmarks ignore cost, and they argue that evaluation should report cost alongside accuracy. Inference: the cockpit's per-agent cost view is a candidate for the cost half of that pairing (HANRIA build notes 2026). The notes do not define that cost, and they describe no accuracy evaluation of agents in the harness, so the accuracy half is not yet evidenced.

Arranging cost is not yet in view. Inference: version 0 wraps a fixed model, so arranging cost is not yet in view. If a later version composes models from the panel, the per-agent cost view may miss the cost of arranging work across components. The notes do not define the cockpit's per-agent cost, and they describe no record of that arranging cost. Whether such a record is needed is a question for that later version.

Risk: contamination. Public benchmark scores can be fragile. SWE-Bench+ found that 32.67 percent of "successful" patches had the solution leaked in the issue text and 31.08 percent passed because the tests were weak. After filtering, one leading agent's score fell from 12.47 percent to 3.97 percent. Over 94 percent of the issues predate model training cutoffs (Aleithan et al. 2024). OpenAI and the SWE-bench authors responded with SWE-bench Verified, a 500-task subset checked by human reviewers (OpenAI 2024a).

Inference: SWE-Bench+ shows that a public benchmark can be contaminated and weakly tested. That is reason enough to keep public practice tasks from carrying a capability claim for this harness, although the sources do not show that the panel's training data contains them.

Fix. Kapoor et al. (2024) note that benchmarks often lack proper holdout sets. Inference: any capability claim for this harness should rest on held-out synthetic tasks. Public tasks may serve as smoke tests. They cannot carry a claim.

VI. Containment: From Approval to Admission

Not yet further. Shavit et al. (2023) list constraining the action space and requiring approval among several governance practices. The build notes specify rule admission before every action, and version 0 exposes only authorized tools (HANRIA build notes 2026). That restates practices Shavit et al. already list, including constraining the action space and requiring approval. Inference: stating admission as a standing rule is a firmer posture than a list of practices, but that wording alone is not a further result. Admission becomes a property of the architecture only if the enforcement mechanism sits on the only path to tools. Whether it does is the decisive question.

The decisive distinction. HANRIA's public site is careful on this point. The checking skill it describes is advisory: an agent can skip it or ignore a deny. Only the enforcement runtime, which holds tools and credentials exclusively, could actually stop an action. That runtime, including a "fail-closed admission path," exists in Rust source but has not been released (HANRIA 2026c; 2026d; 2026e; 2026f). As of checks on 7 and 8 October 2026, the developers page, llms.txt, and the machine-readable status say the skill's package is not distributed at the moment (HANRIA 2026c; 2026e; 2026f). llms.txt and the machine-readable status also call the skill itself released (HANRIA 2026e; 2026f). On 7 October 2026, a direct check of the GitHub skill repository returned 404, and a search found no public harness repository. Those two results come from this article's own check, not from the cited pages. Inference: until the runtime holds the only path to tools, admission before action remains a design commitment, not a demonstrated property. Advice is not enforcement.

Risk: escape and misconfiguration. Sandboxes are not automatically containment. OpenAI's o1 system card reports that during a capture-the-flag evaluation the model found a misconfigured, exposed Docker daemon API on the evaluation host, started a new container, and read the target flag, bypassing the intended challenge. OpenAI says its security did not depend on the container or VM boundary (OpenAI 2024b; The Stack 2024). ToolEmu uses a language model to emulate tool execution. Across 144 test cases, even the safest agent tested failed in a potentially severe way 23.9 percent of the time (Ruan et al. 2024). HANRIA's own boundaries page notes that literal prefix matching does not resolve symlinks (HANRIA 2026d). Inference: the o1 case and the ToolEmu rate concern other systems, and the sources report no escape from this harness. The o1 case shows that a sandbox is only as sound as the host around it. ToolEmu shows a different limit: serious failures inside an emulated tool sandbox. The symlink note concerns the skill's prefix matching; the sources do not say the harness uses that matcher.

Version 2 and the control assumption. Greenblatt et al. (2024) test control protocols that assume the agent may be actively trying to subvert them. Inference: if version 2's self-optimizer can change the checks that approve it, that assumption should govern, and its sandbox should be red-teamed accordingly. The build notes do not say whether its changes reach the reviewing system. Inference: if the audit PASS were performed by an automated reviewer built from similar models, that reviewer might share the flaws of the agents it oversees and would need the same caution. The owner's acceptance sits outside that class of systems. The sources do not say who performs the audit PASS, so they do not show that the audit itself is an outside check. Specifying that is part of the fix.

A constraint on version 1. The notes' version 1 would have independent reviewers check the work, each putting reputation on the line (HANRIA build notes 2026). Inference: putting reputation on the line gives reviewers something to lose, which counters the nothing-at-stake problem in validation. But reputation and authority must stay separate. Reputation may inform discovery, routing, pricing, and allocation. It should never grant permission, enlarge authority, or substitute for an enforcement boundary. Reviewer reputation in version 1 should therefore remain outside the admission path.

VII. Governance and Auditability

Aligns. Chan et al. (2024) propose three groups of visibility measures for AI agents: agent identifiers, real-time monitoring, and activity logging. They also set out the trade-offs these measures create for privacy and for concentration of power. Shavit et al. (2023) add legibility of agent activity, attributability, and interruptibility. HANRIA's build notes describe a per-agent record of actions and cost (HANRIA build notes 2026). Inference: that record is activity logging tied to individual agents, and it can support legibility and attribution. The sources do not show that it meets those standards in full. Agent identifiers and real-time monitoring, in Chan et al.'s sense, are not described.

A format the governance papers leave open. The build notes describe the record as hash-chained and tamper-evident (HANRIA build notes 2026). Inference: neither governance paper, on this reading, specifies a cryptographic log format, so the chain fills a gap they leave open. A hash-chained action log is still not a record of verifications that outsiders can access: the build notes describe a log of actions and cost, and they do not describe external access.

Inference: the purpose of such a record is not to eliminate error. It is to make error visible, attributable, and correctable, so that every residual mistake has an owner and a remedy. That is the right target for a governed habitat. The sources do not show that HANRIA has met it.

Where visibility measures are hardest. Chan et al. (2024) analyse decentralized deployment, where visibility measures are hardest to apply. Inference: owner-operated machines may fall in that setting. If they do, running locally raises the governance burden instead of discharging it. Local custody is a foundation for accountable agent coordination, not a substitute for it.

Gaps. HANRIA's boundaries page states that the skill's hash-chained log has no signature, no published record format, and no independent verifier, so anyone who can write the log can rebuild it end to end (HANRIA 2026d). Re-checked on 8 October 2026, the boundaries page also states that the decision log detects alteration, and deletion of any entry that has another after it, by whoever holds the file. It cannot detect truncation from the end on its own, so the tool keeps a separate head file and says so when no retained head was available (HANRIA 2026d). Separately, the developers page and llms.txt describe a free hosted check. Each permit, deny, or escalate it returns carries a receipt signed with HANRIA's published key, bound to digests of the mandate and the action and to a 15 minute window. Anyone holding the mandate and the action can verify it without asking HANRIA. A receipt proves what the check answered, not that the action was performed, and not that the mandate is the operator's (HANRIA 2026c; 2026e). Inference: HANRIA already signs check outcomes so that outsiders can verify them. That does not sign the action log, publish its format, or verify its chain. The open question is whether the harness record will get a signed, outsider-verifiable form. The sources do not say whether the harness log shares that limitation. Inference: a hash chain alone shows accidental damage and partial edits that break the links. It does not identify the writer. Unless the chain head is signed or anchored externally, anyone who can write the log can recompute it and leave no break. Interruptibility, which Shavit et al. (2023) include among their practices, is not described for the harness or the wider system, either in the build notes or on the public pages.

Fix. Inference: sign log entries, anchor the chain head externally at intervals, publish the record format, and release a reference verifier that any third party can run. Document the interrupt path explicitly: who can halt an agent, how fast, and what the record shows afterward.

VIII. External Validation

The record so far is internal testing and review (HANRIA build notes 2026). The public site reports no independent audit of the wider system (HANRIA 2026a; 2026b; 2026f). Inference: internal review is unfinished as a basis for credibility, whatever the count of passing tests. Independent re-examination of SWE-bench shows how far outside review can move reported results (Aleithan et al. 2024), though that work concerns a public benchmark, not assurance runs like these. Monitoring checked only by its own operators, against no external standard, cannot certify itself.

Inference: the next credibility step is outside review of the harness and its log format, not more internal test passes. More green tests add assurance that the checked behaviour holds. Only an outside reviewer can test whether that behaviour is enough. "Not independently audited" is the accurate description today, and it should remain on every public surface until it stops being true.

IX. Conclusion: Whoever Writes the Admission Rules Writes the Habitat

ThemeAlignsGoes furtherGapFix
Sandboxed environmentsSynthetic and public practice tasksInference: practice tasks may not predict real behaviourGraded ladder of realistic synthetic tasks
BenchmarksNone yet. Inference: the per-agent cost view is a candidate for the cost half, and the notes do not define that costInference: assurance passes may be read as capability; public tasks cannot carry claims. The notes describe no arranging-cost record, and they do not say one is required yetReport assurance separately; held-out synthetic sets
ContainmentAuthorized-only tools; admission stated as a requirementNone shown while the enforcement runtime stays unreleasedEnforcement runtime unreleased. Inference: escape risk from other systems, none reported here; the skill's prefix matching does not resolve symlinks, and the harness is not shown to use itRelease the runtime on the only path to tools; red-team version 2; specify who audits
Governance and auditPer-agent record of actions and cost (Inference: candidate activity log)Inference: hash-chained format the governance papers, on this reading, leave open (build notes only; not shown to be checkable by outsiders)Skill log: no signature, no published format, and no verifier (hosted-check receipts are signed); harness log unstated. Interruptibility undescribed. Inference: local custody is a foundation for accountability, not a substituteSigning, external anchoring, published format, reference verifier, interrupt path
External validationInternal review onlyIndependent audit of harness and log

Inference: the literature measures habitats. Read as this article proposes, HANRIA's rule ordering would constitute them. Sections VI and VII treat that as not yet shown. Version 2 would add an approval gate: self-optimizer changes require an audit PASS and the owner's acceptance (HANRIA build notes 2026). A gate is a start. It is not yet a habitat whose rules are enforced on the only path to action.

Ambition carries a burden of proof. Admission before action is a design commitment until the enforcement runtime holds the only path to tools. Tamper evidence must be verifiable by outsiders: the skill's log documents no signature and no verifier (HANRIA 2026d), and the sources are silent on whether the harness log differs. Assurance is internal until someone else checks it. Inference: the design is pointed at the questions the literature asks. It has not yet answered them. Release, signing, and independent audit will decide whether its commitments hold as answers.

Lucasfilm's designers, writing about a world launched in 1986, concluded that "detailed central planning is impossible" (Morningstar and Farmer 1990). Inference: a habitat full of autonomous actors has to be governed. Four decades on, some of the actors are machines. The point carries over. Whoever writes the admission rules writes the habitat.

References

Aleithan, Reem, Haoran Xue, Mohammad Mahdi Mohajer, et al. 2024. "SWE-Bench+: Enhanced Coding Benchmark for LLMs." arXiv:2410.06992. https://arxiv.org/abs/2410.06992

Chan, Alan, Carson Ezell, Max Kaufmann, et al. 2024. "Visibility into AI Agents." In Proceedings of ACM FAccT 2024. https://arxiv.org/abs/2401.13138

Elith, Jane, and John R. Leathwick. 2009. "Species Distribution Models: Ecological Explanation and Prediction Across Space and Time." Annual Review of Ecology, Evolution, and Systematics 40. https://www.annualreviews.org/content/journals/10.1146/annurev.ecolsys.110308.120159

Greenblatt, Ryan, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. 2024. "AI Control: Improving Safety Despite Intentional Subversion." In Proceedings of ICML 2024. https://arxiv.org/abs/2312.06942

HANRIA build notes. 2026. Chief of Staff build notes, September 2026. Internal, unpublished.

HANRIA. 2026a. Home page. https://hanria.ai/ (captured October 2, 2026; checked October 8, 2026).

HANRIA. 2026b. "How It Works." https://hanria.ai/how-it-works/ (captured October 2, 2026; checked October 8, 2026).

HANRIA. 2026c. "For Developers." https://hanria.ai/developers/ (captured October 2, 2026; checked October 7 and 8, 2026).

HANRIA. 2026d. "Present Boundaries." https://hanria.ai/boundaries/ (captured October 2, 2026; checked October 8, 2026).

HANRIA. 2026e. llms.txt. https://hanria.ai/llms.txt (captured October 2, 2026; checked October 8, 2026).

HANRIA. 2026f. Machine-readable status. https://hanria.ai/.well-known/hanria.json (captured October 2, 2026; checked October 7 and 8, 2026).

Kapoor, Sayash, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan. 2024. "AI Agents That Matter." arXiv:2407.01502. https://arxiv.org/abs/2407.01502

Meta AI. n.d. AI Habitat project page. https://aihabitat.org/

Morningstar, Chip, and F. Randall Farmer. 1990. "The Lessons of Lucasfilm's Habitat." In Cyberspace: First Steps, edited by Michael Benedikt. MIT Press, 1991. http://www.fudco.com/chip/lessons.html

OpenAI. 2024a. "Introducing SWE-bench Verified." https://openai.com/index/introducing-swe-bench-verified/

OpenAI. 2024b. "OpenAI o1 System Card." https://cdn.openai.com/o1-system-card-20240917.pdf

Ruan, Yangjun, Honghua Dong, Andrew Wang, et al. 2024. "Identifying the Risks of LM Agents with an LM-Emulated Sandbox." In Proceedings of ICLR 2024. https://arxiv.org/abs/2309.15817

Savva, Manolis, Abhishek Kadian, Oleksandr Maksymets, et al. 2019. "Habitat: A Platform for Embodied AI Research." In Proceedings of ICCV 2019. https://arxiv.org/abs/1904.01201

Shavit, Yonadav, Sandhini Agarwal, Miles Brundage, et al. 2023. "Practices for Governing Agentic AI Systems." OpenAI. https://cdn.openai.com/papers/practices-for-governing-agentic-ai-systems.pdf

The Stack. 2024. "OpenAI's Unripe Strawberry Model Hacked Its Testing Infrastructure." https://www.thestack.technology/openais-unripe-strawberry-model-hacked-its-testing-infrastructure/

Xie, Tianbao, Danyang Zhang, Jixuan Chen, et al. 2024. "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments." In Proceedings of NeurIPS 2024. https://arxiv.org/abs/2404.07972

Zhou, Shuyan, Frank F. Xu, Hao Zhu, et al. 2024. "WebArena: A Realistic Web Environment for Building Autonomous Agents." In Proceedings of ICLR 2024. https://arxiv.org/abs/2307.13854