Open evidence for Eliza AI.
No result without a reproducible trail. This page separates what Eliza can demonstrate today from what ELM HQ intends to build. Raw logs, dataset hashes, versions and failed cases belong beside every score.
A measurable index, with missing abilities left missing.
Preview · coverage gatedThe Eliza Intelligence Index (EII) combines versioned capability tests without pretending that one narrow score is Eliza's IQ. The current observed-domain score is 56.40 / 100, but it covers only 35% of the planned index and only 20% currently comes from a standard external benchmark. Therefore no overall EII score is issued yet.
General reasoning (MMLU-Pro), agentic tool use (GAIA), and multimodal understanding (MMMU) must be run with public raw logs. Embodied robotics also remains unmeasured. The 10-case conversation score is a component diagnostic, not an external standard benchmark or complete deployed-path test.
Standard benchmarks before sweeping claims.
Complete · provisional local judgeEliza's earlier 12-question synthetic-age test remains useful as a deployed-path smoke test, but it is not a substitute for a standard benchmark. A 500-question run on the pinned official cleaned LongMemEval oracle split is complete. The answers and every verdict are downloadable; because qwen3.5:9b acted as both reader and judge, this first score is provisional rather than independently audited.
The required session was retrieved in every answerable case, yet provisional QA accuracy was 61.60%, multi-session accuracy was 39.10%, and false recall on unanswerable questions was 60.00%. Those failures are now explicit improvement targets rather than hidden behind the retrieval score.
Inspect all 500 case recordsA changed module is not automatically a gained skill.
Baseline recorded · no gain claimedThe current record contains installed and expanded modules, but also many duplicate proposals and failures. A public improvement claim now requires the same versioned task suite before and after a change, raw outputs, code/runtime versions and at least seven real elapsed days.
Hardware only counts when hardware completes the task.
Not demonstratedThe Bodyshop and robotics partnership route describe engineering work and invite hardware access. They do not establish locomotion or manipulation performance. Until repeated physical trials exist, the published success rate remains blank rather than inferred from diagrams or simulations.
Recalculate it yourself.
Each completed run publishes a summary and line-by-line raw log with question type, retrieved evidence, answer, verdict, latency, dataset hash and declared models. Anyone can rescore the outputs with a different judge.
The first completed score used the same local reader and judge, so it remains provisional. A second local model is now regrading the fixed published answers to measure judge agreement. That reduces same-model circularity but will remain labelled separately from an external audit.
Guest sessions are intentionally anonymous and best-effort; they are not evidence for durable person-level memory. Persistent history and cross-device continuity are evaluated only for authenticated accounts.