Open research project · Local-first · Human-governed

Hu-Mind AI

An experiment in building AI that does not simply accept its first answer. Hu-Mind asks distinct reasoners to examine a problem independently, exposes their proposals to a bounded challenge, verifies the evidence and permits consequential action only after exact agreement and human approval.

An important distinction: Hu-Mind is not AGI today and is not being presented as one. AGI is a long-term research target that would require preregistered tests, independent replication and external assessment. The present system is an advisory prototype with no authority to edit files, run commands or publish actions.

The question

Can structured disagreement improve AI decisions?

A fluent answer can still be wrong. Two models can also repeat the same hidden assumption. Hu-Mind is designed to make claims, disagreements, evidence and authority visible and testable.

01

Generate alternatives

A creative reasoner proposes useful options, hypotheses and possible actions instead of collapsing immediately onto one answer.

02

Challenge assumptions

A bounded “shadow probe” raises one constructive counter-hypothesis. It is not a personality or a stream of negative thoughts; its purpose is to expose overlooked failure modes.

03

Review independently

Two heterogeneous reasoners assess evidence, contradictions, feasibility and risk before either sees the other’s conclusion.

04

Require accountable consent

A deterministic gate checks exact action agreement, confidence, evidence and vetoes. Agreement between models is still insufficient: a human remains the final authority.

Proposed architecture

From a brief to a governed decision

The popular “left brain / right brain” comparison is only a useful design metaphor. Hu-Mind is not a psychological model and does not claim to reproduce a human brain.

  1. 1
    Human briefPurpose, boundaries and available evidence
  2. 2
    Bounded challengeA constructive counter-hypothesis, adjusted within strict limits
  3. 3
    Independent reasonersDistinct local models review without copying one another
  4. 4
    Evidence and synthesisClaims are linked to sources; disagreement triggers revision or escalation
  5. 5
    Deterministic consensusExact actions, critical vetoes and resource limits are checked in code
  6. 6
    Human authorityConsequential action remains blocked until explicitly approved

Offline sovereignty

The reference system is intended to operate inside a closed environment after signed models and artefacts have been provisioned. Cloud models may be used as research comparators during development, not as the final dependency.

Evidence before scale

More GPUs are not a substitute for a sound experiment. Model fit, throughput, memory, energy and failure behaviour must be measured before asking for larger hardware.

Productive disagreement

The system should not force consensus. If evidence is missing or reviews conflict, the correct result may be to request more information or escalate to a person.

Reversible progress

Experiments remain sandboxed, rate-limited, auditable and capable of rollback. Negative results and abandoned hypotheses are treated as useful research outputs.

What exists now

A working safety and evaluation foundation

The current repository goes beyond a concept document. Its local Python prototype includes the components needed to begin controlled comparisons, while deliberately leaving tool execution disabled.

  • Bounded adaptive shadow-probe generation
  • Structured independent reviews and one redacted critique round
  • Exact-action consensus, confidence thresholds and critical vetoes
  • Content-addressed evidence packs and claim-link verification
  • Hash-chained experiment plans and results
  • A validated 100-task evaluation format and shared scoring
  • Token, latency, energy and peak-memory telemetry fields
  • Provenance-aware memory with correction, retraction and decay
  • Offline provider manifests, mocks and failure-path tests
  • A human approval and action-allowlist boundary
Research roadmap

Progress is controlled by evidence gates

Each stage must answer a specific question before capability, autonomy or hardware is expanded.

Gate 0

Reproducible foundation

Rebuild the offline tests from signed artefacts on a second machine.

Gate 1

Two local reasoners

Run two genuinely distinct models without a runtime internet dependency and fail closed under provider errors.

Gate 2

Falsify the design

Compare the architecture with cost-matched single-model baselines and publish negative results.

Gate 3

Controlled learning

Test memory, transfer, correction, poisoning resistance and rollback on unseen task families.

Gate 4

Bounded local agent

Evaluate long-horizon planning inside an offline simulation with narrow, supervised actions.

Gate 5+

Independent assessment

Only external replication and plural governance could justify moving toward any candidate-AGI claim or shared facility.

Hardware and funding

Funding should buy an experiment, not a promise

The immediate aim is not to purchase the largest available machine. It is to run the smallest credible dual-model experiment, record the bottlenecks, then request the exact compute needed for replication.

Initial guidance emails were sent on 18 September 2026 to NVIDIA Inception, the UKRI AI Research Resource and the University of Bath’s partnerships team. They are exploratory approaches only: no response, endorsement, partnership, compute award or funding commitment has been recorded.

1Verify the SIMBA hardware and signed model manifests
2Measure tokens, latency, memory, energy and thermals
3Run the fixed evaluation and document failures
4Ask partners for a specific node, replication run or research contribution
Open collaboration

How other developers and researchers can help

The most valuable contributions are testable, reviewable and small enough to understand. You do not need expensive hardware to improve the project.

Engineering

Review orchestration, add local-model adapters, improve schema validation, strengthen failure handling or make telemetry reproducible.

Evaluation

Design held-out tasks, single-model baselines, ablations and scoring methods that can disprove the project’s hypotheses.

Safety & security

Threat-model prompt injection, correlated model errors, evidence poisoning, memory attacks, sandbox escape and approval bypass.

Local AI & hardware

Replicate inference on modest hardware, publish model-fit measurements and help create signed, offline deployment manifests.

Research partnerships

Help with independent replication, ethics, governance, red teaming, grant routes or access to time-limited compute.

Documentation

Improve plain-English explanations, diagrams, accessibility, experiment records and contributor onboarding.

A good first contribution

Read the architecture and roadmap, choose one bounded issue, explain the failure or improvement you are targeting, and open an issue before a large change. Include tests with code and never commit credentials, private data or unaudited model output.

Build the evidence with us

Challenge the idea. Improve the tests. Help us find out what works.

Hu-Mind will be useful only if its claims survive careful criticism and independent reproduction.