Research & development

Aegis Humaniform.

Researching AI that can remember, investigate, verify, act, and adapt across long-running work.

This is an active public-safe research record. It explains the direction, the evidence currently available, and the questions that still need proof.

What problem is Humaniform researching?

Current AI can be capable within one session and still lose continuity: important decisions, source evidence, corrections, and the reason a tool action was taken.

Forget decisions between interactions

Lose useful evidence in large contexts

Repeat mistakes that were already corrected

Act without a clear record of why

What is actively being researched?

Humaniform is an integrated research architecture, not a bundle of disconnected features. Each area below has a distinct technical question and a practical outcome to test.

01

Evidence-grounded cognition

Keep important claims connected to their origin with source-linked memory, freshness tracking, contradiction handling, and separate confidence and evidence estimates.

An AI system that can explain what supports a conclusion, what is missing, and what would change it.

02

Persistent context and memory

Investigate layered memory for working context, decisions, source material, project history, semantic knowledge, skills, and scoped adaptation.

Continuity across long software projects, changing codebases, research programs, and operational workflows.

03

Passive and active comprehension

Observe entities, events, changes, anomalies, and unresolved contradictions before a question arrives, then investigate deliberately when a goal or uncertainty appears.

A system that is prepared to reason about a living environment rather than only reacting to the latest prompt.

04

Adaptive computation

Explore fast paths for familiar work and deeper computation for ambiguity, novelty, risk, or long-horizon planning.

More useful work per unit of compute without treating every input as equally difficult.

05

Predictive world models

Represent entities, state, events, actions, dependencies, and possible futures so the system can compare outcomes before acting.

Better planning, side-effect prediction, recovery after failure, and investigation of complex systems.

06

Governed agents

Give tool-using systems explicit identity, capabilities, data scope, risk budgets, evidence requirements, reversibility, and approval thresholds.

Increasingly capable agents that preserve human control, traceability, and clear boundaries.

07

Continual adaptation

Separate fast episode state, scoped reversible changes, and persistent updates accepted only after verification, replay, and retention testing.

Learning from experience without treating every interaction as permanent truth.

08

Recursive investigation

Treat external context as an environment to investigate through metadata, bounded retrieval, subproblems, specialists, deterministic tools, and source-backed synthesis.

Effective operation over large repositories, document collections, histories, and research corpora while active memory stays bounded.

09

Learning from failure

Turn wrong sources, failed patches, missed contradictions, unsupported actions, and repeated strategies into structured replay and correction signals.

A persistent system that makes fewer repeated mistakes, not simply one that stores more conversations.

10

Efficient model development

Study waste across low-value examples, redundant supervision, uniform computation, long-context activation cost, and unnecessary updates.

A research path toward more efficient training and post-training, evaluated with complete cost accounting.

How would the system work?

The current model is an integrated stack. It keeps context and evidence close to the work while letting the system investigate uncertainty and act through controlled tools.

  1. 01

    Observe

    Streams, documents, code, interfaces, and tools

  2. 02

    Comprehend

    Entities, events, goals, uncertainty, and predictions

  3. 03

    Remember

    Exact evidence, important events, and project state

  4. 04

    Investigate

    Retrieval, recursion, specialists, and tools

  5. 05

    Reason

    Alternatives, causality, world models, and planning

  6. 06

    Verify

    Evidence, tests, contradictions, and permissions

  7. 07

    Act

    Controlled tools, software changes, and workflows

  8. 08

    Learn

    Replay, scoped adaptation, and validated consolidation

System layers

Humaniform Core

General reasoning, prediction, uncertainty, and adaptive computation.

Aegis Evidence

Source-linked memory, provenance, verification, and claim support.

Continuum

Persistent context, project memory, and long-running state.

AURORA

Retrofitting Humaniform capabilities onto existing models.

FORGE

Research into more efficient training, post-training, and continual improvement.

What evidence is available today?

The following describes controlled prototype evidence and its limits. It is not a claim of broad general intelligence, universal reliability, or an independently validated production result.

Causal grounding

Source-grounded answer behavior survived in the learned Humaniform path and collapsed in no-anchor and transformer controls.

Reproduced controlled prototype result

Exact source linkage

The model learned to produce source spans and citations instead of answer text alone.

Demonstrated on controlled grounding fixtures

Evidence recall architecture

A separate high-recall source bank raised supported source exposure and supported answering to 100% on the current fixture.

Controlled-fixture architectural validation

Causal ablations

Removing, masking, shuffling, or corrupting evidence caused target grounded behavior to collapse or degrade.

Direct causal evidence

Learned Humaniform

Joint answer + citation

0.3984

Source-span accuracy

0.6680

Confident-wrong rate

0.0000

Directly documented · causally supported

No-anchor control

Joint answer + citation

0.0000

Controlled comparison

Plain transformer

Joint answer + citation

0.0000

Controlled comparison

Correct source exposed top-3

1.0000

Correct source used by support

1.0000

Supported-answer accuracy

1.0000

Supported over-refusal

0.0000

Exact learned-anchor presence

0.3594

What still needs proof?

Matched baselines

Compare optimized alternatives with the same hardware, data opportunity, time budget, tools, context, and evaluation suite.

Protected capabilities

Reject gains that materially damage language quality, reasoning, coding, grounding, calibration, safety, or generalization.

Causal ablations

Remove or disable components to test whether they actually cause the claimed benefit.

Full cost accounting

Include training, processing, retrieval, verification, replay, storage, communication, failed runs, and runtime overhead.

Fresh-task testing

Evaluate on new repositories, altered environments, hidden future questions, counterfactual changes, and adversarial inputs.

Rollback conditions

Pause or remove a subsystem when overhead, instability, or capability regressions outweigh its value.

Explicitly outside the current claim boundary

Broad general-intelligence superiority

Human-equivalent comprehension

Safe autonomous self-improvement

Unlimited or zero-rot context

Universal agent reliability

Guaranteed training-compute reductions

Fully autonomous software development

Scientific or industrial autonomy

Consciousness or AGI

Where does the work go next?

  1. Stage 1

    Grounding and evidence

    Source-linked memory, exact pointers, evidence selection, contradiction handling, and grounded action control.

  2. Stage 2

    Throughput and efficiency

    Fast paths, fixed-shape routing, fused evidence operations, useful-work measurement, and matched compute accounting.

  3. Stage 3

    Passive comprehension

    Event and entity tracking, future-state prediction, deferred-query memory, and proof-preserving compression.

  4. Stage 4

    Persistent context

    Layered memory, large external context, bounded active state, exact source fallback, and context-integrity evaluation.

  5. Stage 5

    Governed agents

    Tools, software workflows, permissions, verification, recovery, audit, and rollback.

  6. Stage 6

    Continual adaptation

    Temporary neural memory, scoped adapters, verified local learning, retention, poisoning resistance, and safe consolidation.

  7. Stage 7

    Native Humaniform models

    Architecture scaling, specialist computation, efficient training, multimodal capability, and long-running deployment.

Near-term validation priorities

01

Restore and audit the later proof package

Confirm the owner-reported C2h, Gate D, and Gate E results from raw artifacts.

02

Invalid-evidence rejection

Test wrong entity, wrong field, wrong source, stale, contradictory, ambiguous, and corrupted bindings.

03

Complete support envelopes

Raise complete human-auditable source-span recall from the reported zero baseline.

04

Exact top-source selection

Improve exact top-one selection while preserving high recall.

05

Fixed-quality efficiency

Measure target quality, grounded work per second, latency, FLOPs, energy, cost, and memory.

Make a hard question testable.

If you are researching a serious AI problem, governed agent capability, or specialised business model, start with the evidence you need and the risk you need to control.

Start a research conversation