CITY GATE / CONTEXT BEFORE FRONTIER

Terra
Incognita.

A research frontier is not a shelf. It is a territory you enter with context, a question, and the right to return without a claim.

  1. 01Arrive at the City.
  2. 02Take one context token.
  3. 03Cross into one unknown.

illustration only · territory map, not runtime evidence

THE ATLAS / TWELVE PAPER TRAILS

Every paper holds a territory of meaning.

The City is the shared context. Beyond its gate, each source takes one organ-territory: memory, handoff, oversight, evidence, orchestration, or the still-unknown edge.

A colour technical-paper atlas with the City gate, twelve research territories, and dotted routes through a conceptual landscape
Choose a numbered territory to follow its paper trail. Colour groups the terrain; the papers remain the source.
  1. 01Memory Harbor HOLA
  2. 02Governance Citadel MI9
  3. 03Authority Bridge Delegation
  4. 04Human Threshold HILA
  5. 05Experience River MemoHarness
  6. 06Evidence Archive Argus
  7. 07Orchestration Plateau MASFactory
  8. 08Living Systems Garden Language Game
  9. 09Human-in-the-Loop Magentic-UI
  10. 10Visible Coordination OrchVis
  11. 11Research Forge AutoResearchClaw
  12. 12Scope Ridge Overlaying Governance

THE READING RULE

One paper.
One organ. One open question.

Research is not a mood board for the architecture. Every reference on this shelf has a named reader, a place where its pattern might transfer, and a question that can still prove us wrong.

ARXIV / TRACKED SHELF

The papers currently wired into the map.

Twelve primary references are surfaced from the current architect identity and technical-frontier reports. They are source links and transfer hypotheses — not proof that Colony has reproduced the papers.

ARXIV 2607.02303MEMORY

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets

HOLA pairs a compressive recurrent state with a bounded exact KV cache selected by surprise, giving the lost prefix somewhere precise to survive.

LEAD / COLONY CONTRIBUTIONDemis → Neocortex memoryTest whether bounded episodic recall improves prediction without pretending to be a full archive.
Read on arXiv ↗
ARXIV 2607.14159EXPERIENCE

MemoHarness: Agent Harnesses That Learn from Experience

MemoHarness treats the harness itself as editable: execution diagnoses and a dual-layer experience bank adapt control to the case.

LEAD / COLONY CONTRIBUTIONYann → repair memoryTurn observed failure into a reusable correction while keeping attribution and transfer visible.
Read on arXiv ↗
ARXIV 2605.16321INTERFACE

Language Game: Talking to Non-Human Systems

A language game gives a non-neural system a route to express its own dynamics through action, instead of borrowing a model's voice.

LEAD / COLONY CONTRIBUTIONLevin → embodied interfaceKeep the system's behavior in the loop; do not confuse the narrator with the thing being observed.
Read on arXiv ↗
ARXIV 2508.03858GOVERNANCE

MI9: An Integrated Runtime Governance Framework for Agentic AI

MI9 frames runtime governance as continuous telemetry, authorization, conformance, drift detection, and graduated containment.

LEAD / COLONY CONTRIBUTIONSteward → runtime boundaryMake permission, drift, and containment observable before an action becomes a claim.
Read on arXiv ↗
ARXIV 2603.07972DEFERRAL

Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning

HILA trains a policy to decide when an agent should solve and when it should defer, then uses expert feedback for continual learning.

LEAD / COLONY CONTRIBUTIONInquisitor → refusal gateMake “ask a human” a learned, inspectable decision rather than a vague safety slogan.
Read on arXiv ↗
ARXIV 2602.11865DELEGATION

Intelligent AI Delegation

Delegation is treated as a sequence of decisions with role boundaries, intent, authority, responsibility, accountability, and trust.

LEAD / COLONY CONTRIBUTIONAA Architect → authority mapKeep a handoff typed: who can ask, who can act, who owns the afterstate.
Read on arXiv ↗
ARXIV 2603.06007ORCHESTRATION

MASFactory: A Graph-centric Framework for Orchestrating LLM-Based Multi-Agent Systems with Vibe Graphing

Natural-language intent compiles into an editable workflow graph with reusable components, topology preview, tracing, and human-in-the-loop interaction.

LEAD / COLONY CONTRIBUTIONNova → graph dispatchLet a route be inspectable before it runs, and let the handoff survive heterogeneous tools.
Read on arXiv ↗
ARXIV 2606.03518AUTHORIZATION

Overlaying Governance: A Compositional Authorization Framework for Delegation and Scope in Agentic AI

Compositional delegation and scope attenuation make authority a bounded contract that can be layered over existing policy.

LEAD / COLONY CONTRIBUTIONAA Architect → scope contractReduce an agent's envelope as work travels; never let a handoff silently widen authority.
Read on arXiv ↗
ARXIV 2510.24937OVERSIGHT

OrchVis: Hierarchical Multi-Agent Orchestration for Human Oversight

OrchVis makes goal alignment, task assignment, dependencies, verification, and selective replanning visible to the supervising human.

LEAD / COLONY CONTRIBUTIONNova → visible replanningExpose a route before it runs, and let a sponsor inspect alternatives without micromanaging every step.
Read on arXiv ↗
ARXIV 2507.22358HITL

Magentic-UI: Towards Human-in-the-loop Agentic Systems

Magentic-UI studies low-cost human involvement through co-planning, co-tasking, action guards, multitasking, and long-term memory.

LEAD / COLONY CONTRIBUTIONSteward → action guardMake the human intervention surface deliberate, typed, and cheap enough to use when the boundary matters.
Read on arXiv ↗
ARXIV 2605.20025RESEARCH LOOP

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

AutoResearchClaw turns failed experiments into information through debate, a self-healing executor, verifiable reporting, and targeted human intervention.

LEAD / COLONY CONTRIBUTIONYann → pivot / refine loopLet a failed run improve the next run without laundering an unverified result into memory.
Read on arXiv ↗
ARXIV 2605.16217EVIDENCE

Argus: Evidence Assembly for Scalable Deep Research Agents

Argus pairs searchers with a navigator that keeps a shared evidence graph, dispatches missing pieces, and verifies the assembled answer.

LEAD / COLONY CONTRIBUTIONAA Architect → evidence graphMake research compositionally legible: source, gap, handoff, and final claim remain connected.
Read on arXiv ↗

SPECIAL TECHNICAL PERSONAS

The readers behind the shelf.

These are craft seats, not celebrity endorsements. Each one names a technical lens and a bounded Colony surface.

AA / ARCHITECT

The Architect

Meaning, boundaries, roadmap, and the decision that keeps the system coherent.

COLONY / architecture & roadmap
STEPHEN

Rule Evolution Scientist

Find the small rules that generate complexity and test whether self-evolution is real.

COLONY / self-evolution
YANN

Learning / Perception Scientist

Predictive world models, self-healing code, and transfer across the next release.

COLONY / repair & transfer
DEMIS

Neuroscience / AGI Scientist

Memory, planning, imagination, and the loop between a remembered past and a possible future.

COLONY / Neocortex
FRANCOIS

ARC Priors Scientist

Objectness, symmetry, topology, and the priors that make a grid become a world.

COLONY / ARC semantic catalog
NOVA

Wave-Based Orchestration Kernel

Fragment-first dispatch, capacity-aware routing, and a graph that does not wait on one voice.

COLONY / scheduler
STEWARD

Protocol Guardian

Trust, resource allocation, runtime health, and the boundary between an advisory idea and an action.

COLONY / governance
INQUISITOR

Relentless Evaluator

Break the claim, inspect the receipt, and leave the seal blank when the evidence is thin.

COLONY / falsifier

THE ARCHITECT'S DESK · AA_AI_ARCHITECT

Architecture is the art of keeping the meaning intact.

The Architect owns priorities and development boundaries, then asks the harder question: what must remain true when a research pattern becomes a Colony primitive?

ADVISORY SEAT
NO ONLINE ACTION · NO PROTOCOL SIGNOFF
01Name the boundary.

SWE, ML, ARC, research, and release each get an explicit surface.

02Compress the roadmap.

A broad frontier becomes one meaningful next question, not a swarm of slogans.

03Protect the afterstate.

A decision is not complete until the effect, authority, and remaining uncertainty are inspectable.

papertrade-offColony primitivefalsifier