ELM HQ

Open evidence for Eliza AI.

No result without a reproducible trail. This page separates what Eliza can demonstrate today from what ELM HQ intends to build. Raw logs, dataset hashes, versions and failed cases belong beside every score.

A measurable index, with missing abilities left missing.

Preview · coverage gated

The Eliza Intelligence Index (EII) combines versioned capability tests without pretending that one narrow score is Eliza's IQ. The current observed-domain score is 56.40 / 100, but it covers only 35% of the planned index and only 20% currently comes from a standard external benchmark. Therefore no overall EII score is issued yet.

OVERALL EIINot issued70% standard coverage required
OBSERVED-DOMAIN SCORE56.40measured domains only
OBSERVED COVERAGE35%memory + conversation diagnostic
STANDARD COVERAGE20%LongMemEval
Three critical domains remain unmeasured.

General reasoning (MMLU-Pro), agentic tool use (GAIA), and multimodal understanding (MMMU) must be run with public raw logs. Embodied robotics also remains unmeasured. The 10-case conversation score is a component diagnostic, not an external standard benchmark or complete deployed-path test.

Standard benchmarks before sweeping claims.

Complete · provisional local judge

Eliza's earlier 12-question synthetic-age test remains useful as a deployed-path smoke test, but it is not a substitute for a standard benchmark. A 500-question run on the pinned official cleaned LongMemEval oracle split is complete. The answers and every verdict are downloadable; because qwen3.5:9b acted as both reader and judge, this first score is provisional rather than independently audited.

PROVISIONAL QA ACCURACY61.60%500-question oracle run
TEMPORAL REASONING48.87%65 / 133 judged correct
FALSE RECALL ON ABSTENTION60.00%important open failure
RETRIEVAL SESSION RECALL@K100%retrieval did not guarantee a correct answer
Retrieval is stronger than answer use.

The required session was retrieved in every answerable case, yet provisional QA accuracy was 61.60%, multi-session accuracy was 39.10%, and false recall on unanswerable questions was 60.00%. Those failures are now explicit improvement targets rather than hidden behind the retrieval score.

Inspect all 500 case records

A changed module is not automatically a gained skill.

Baseline recorded · no gain claimed

The current record contains installed and expanded modules, but also many duplicate proposals and failures. A public improvement claim now requires the same versioned task suite before and after a change, raw outputs, code/runtime versions and at least seven real elapsed days.

EXPERIMENT RECORDS655–20 September 2026
INSTALLED3activity, not proof of gain
DUPLICATE44retained, not hidden
FIXED BASELINE16 / 16narrow safeguards suite
MEASURED GAINNone yetsecond run after 7 real days required

Hardware only counts when hardware completes the task.

Not demonstrated

The Bodyshop and robotics partnership route describe engineering work and invite hardware access. They do not establish locomotion or manipulation performance. Until repeated physical trials exist, the published success rate remains blank rather than inferred from diagrams or simulations.

PHYSICAL TRIALS0standard protocol
SUCCESS RATE—no denominator yet
SAFETY STOPS—reported with trials
HUMAN INTERVENTIONS—reported, never omitted

Recalculate it yourself.

Each completed run publishes a summary and line-by-line raw log with question type, retrieved evidence, answer, verdict, latency, dataset hash and declared models. Anyone can rescore the outputs with a different judge.

The first completed score used the same local reader and judge, so it remains provisional. A second local model is now regrading the fixed published answers to measure judge agreement. That reduces same-model circularity but will remain labelled separately from an external audit.

Guest sessions are intentionally anonymous and best-effort; they are not evidence for durable person-level memory. Persistent history and cross-device continuity are evaluated only for authenticated accounts.