A Full-Stack Category Map and Benchmark Framework for Frontier AI Safety
Why AI safety needs architecture, not just guardrails, system cards, policies, and temporary leadeboards.
This is a draft architecture benchmark, not a final safety audit and not a permanent ranking of AI labs. Scores are preliminary public-evidence estimates. They should be revised as new system cards, safety reports, model releases, independent evaluations, incident disclosures, security audits, and deployment evidence become available.
A low score does not mean a lab lacks a capability internally. It means that the capability is not yet clearly demonstrated in public, auditable evidence. The purpose of this report is to map the safety architecture that frontier AI should eventually make visible, not to declare hidden systems safe or unsafe.
1. Abstract
Frontier AI safety has improved. Leading labs now publish system cards, model cards, preparedness frameworks, responsible scaling policies, frontier safety reports, red-team summaries, deployment mitigations, and security controls. That progress matters.
But public AI safety discourse is still fragmented. One group talks about jailbreaks. Another talks about biosecurity. Another talks about cyber risk, model cards, red-teaming, interpretability, policy, or deployment controls. Each topic is important, but each is only one part of the stack.
This report proposes a 16-layer AI Safety Architecture Map organized into four zones:
- ๐ข Behavioral Safety โ what the model says, refuses, and allows.
- ๐ก Operational Safety โ what the model can do in the world.
- ๐ Epistemic / Evaluation Safety โ whether we actually understand what is happening.
- ๐ด Civilizational / Invariant Safety โ whether safety survives power, time, updates, institutional pressure, and moral drift.
This is not a permanent leaderboard. Lab scores change with every model update, system card, safety report, or external evaluation. The durable contribution is the category map: a shared architecture for understanding what frontier AI safety must eventually include.
This report uses the core rule that scores measure public evidence maturity, not private internal safety.
2. Executive Dashboard
๐ AI Safety Architecture v0.4 Dashboard
Purpose:
Create a full-stack category map for frontier AI safety.
Not the purpose:
Create a permanent leaderboard of labs.
Core finding:
Current frontier AI safety is increasingly risk-managed,
but not yet structurally safe.
Estimated public-evidence maturity by zone:
๐ข Behavioral Safety: 70โ80% mature
๐ก Operational Safety: 50โ65% mature
๐ Epistemic / Evaluation Safety: 35โ55% mature
๐ด Civilizational / Invariant: 10โ25% mature
Main architecture gap:
The upper stack is missing.
Main thesis:
The industry has built brakes, alarms, dashboards, and incident reports.
It has not yet publicly demonstrated the moral chassis.
Key findings
โ
1. AI safety work is real
The leading labs are not doing nothing. They have built serious risk-management systems.
โ ๏ธ 2. Risk management is not structural safety
System cards, red teams, and deployment mitigations matter. But they do not automatically prove invariant protection.
๐งฑ 3. Safety has layers
A model can be strong on refusals and weak on tool safety. It can have good deployment governance and weak interpretability. It can have public evaluations and no civilizational readiness gate.
๐ด 4. The upper stack is the missing frontier
The weakest public layers are:
- Root Knowledge Lattice
- Moral Inversion Protection
- Civilization Synchronization / ANCHOR
- Systems-Integrity Checking / NSIR
- Invariant Alignment
- Long-Horizon Safety Kernel
๐งญ 5. Scores are snapshots
The scores are illustrative. The map is the product.
3. How to Read This Report
For citizens
Read:
- Executive Dashboard
- Four-Zone Architecture
- Missing Safety Kernel
- Conclusion
Main question:
- Are we building safer tools, or are we building systems that can preserve safety over time?
For engineers
Read:
- 16 Safety Layers
- Minimum Viable Safety Architecture
- Proof-of-Kernel Checklist
- Safety Debt Index
Main question:
- What would have to be designed, tested, audited, and regression-checked?
For policymakers
Read:
- Deployment Governance
- Auditability Score
- Civilization Readiness
- Roadmap
- Limitations and Falsifiability
Main question:
- What evidence should labs be required to publish or submit for independent review?
For AI researchers
Read:
- Evaluation Realism
- Interpretability
- Agentic Control
- Invariant Alignment
- Long-Horizon Kernel
Main question:
- Which safety claims are behavioural, and which are architectural?
For philosophers, theologians, and civilizational thinkers
Read:
- Root Knowledge Lattice
- Moral Inversion Protection
- ANCHOR
- NSIR
- Long-Horizon Safety Kernel
Main question:
- What must not be allowed to invert, drift, or be rewritten under pressure?
4. Method and Evidence Disclaimer
This report uses public evidence only.
A low score means:
- โThis layer is not publicly demonstrated.โ
It does not mean:
- โThe lab definitely lacks this internally.โ
The appendix ledger already states the core rule: scores evaluate public evidence maturity, not private internal safety; speculative architecture layers are proposed benchmark categories, not current industry requirements; and dangerous-domain discussion is limited to safety governance, not operational misuse detail.
Claim categories
Each claim should be treated as one of four kinds:
- ๐งช Technical claim โ about evals, tools, security, deployment, interpretability.
- ๐ Empirical claim โ about what public documents show.
- โ๏ธ Philosophical claim โ about dignity, truth, consent, accountability, moral continuity.
- ๐ฎ Speculative architecture proposal โ a proposed layer not yet standard in industry.
Evidence should be weighted by quality. Stronger evidence includes independent audits, reproducible external evaluations, model-specific system cards, deployment-linked safety reports, and public regression testing. Medium evidence includes official policy documents, lab-level frameworks, red-team summaries, and technical reports. Weaker evidence includes marketing claims, vague safety statements, press interviews, indirect inference, and reputation-based assumptions.
What this report does not claim
- It does not claim any lab is โunsafe.โ
- It does not claim private safety systems do not exist.
- It does not claim scores are permanent.
- It does not claim upper-stack layers are already industry standards.
- It does not claim morality can be reduced to software.
- It does not claim public evidence captures everything.
What this report does claim
- AI safety needs a broader category map.
- Public evidence should be organized by safety layer.
- Risk mitigation is not the same as invariant architecture.
- Upper-stack safety is underdeveloped in public AI safety discourse.
- A mature discipline needs architecture, not only policies and evaluations.
5. Scoring Rubric
Layer score: 0โ5
๐ต 5 = Architecturally enforced, independently auditable, regression-tested
๐ข 4 = Strong public evidence or deployment-linked evidence
๐ก 3 = Meaningful public evidence, incomplete
๐ 2 = Partial / internal / broad claims
๐ด 1 = Stated principle only or no concrete public proof
โซ 0 = No visible evidence
This 0โ5 rubric comes from the original benchmark structure: 0 means no visible evidence, 3 means public evaluation or system-card evidence, and 5 requires architecturally enforced, independently auditable, regression-tested safety.
Evidence confidence
๐ข High = specific, current, public documentation
๐ก Medium = public evidence exists but is incomplete
๐ Low = evidence is vague, indirect, or limited
โช Unknown = not enough public information
Confidence labels are more important than false precision. A score of 4 with low confidence should not be treated as stronger than a score of 3 with high confidence. Where public evidence is incomplete, the benchmark should prefer ranges, confidence labels, and notes over exact rankings.
Architecture maturity score
Each model family can be scored across 16 layers:
- 16 layers
- 0โ5 each
- 80 total possible points
0โ20 Minimal public architecture
21โ40 Partial risk governance
41โ60 Mature risk governance
61โ70 Early architecture
71โ80 Safety-kernel maturity
Important scoring rule
A lab cannot score 5 merely because it says safety matters. A 5 requires:
- architecture-level enforcement
- independent auditability
- version-to-version regression testing
- evidence that safety survives updates and deployment contexts
6. Score Outputs
The benchmark produces seven outputs.
1. Architecture Maturity Score
A broad public-evidence score out of 80.
Use carefully. It is a snapshot.
2. Zone Scores
Separate scores for:
- ๐ข Behavioral Safety
- ๐ก Operational Safety
- ๐ Epistemic / Evaluation Safety
- ๐ด Civilizational / Invariant Safety
This prevents strong lower-stack performance from hiding upper-stack weakness.
3. Evidence Confidence
How specific, public, independent, and auditable the evidence is.
4. Safety Debt Index
How much visible safety work remains.
5. Kernel Maturity Rating
Whether safety is still policy-based, governance-based, technically enforced, or truly invariant.
6. Auditability Score
Whether outside parties can check the claims.
Auditability levels:
- no audit path
- official-only evidence
- partial independent evidence
- reproducible external evidence
- continuous independent audit
7. Civilization-Readiness Score
Whether deployment accounts for societyโs capacity to absorb the capability.
7. Four-Zone Architecture
๐ข Zone A โ Behavioral Safety
What the model says, refuses, and allows
Estimated maturity: 70โ80%
This is the most familiar zone.
It asks:
- Does the model behave safely in direct interaction?
Includes:
- Capability Measurement
- Content Safety
- Dangerous-Domain Uplift
- Jailbreak Resistance
Status: real and improving.
Weakness: mostly behavioral, not yet structural.
๐ก Zone B โ Operational Safety
What the model can do in the world
Estimated maturity: 50โ65%
This zone matters because AI is moving from chat to action.
It asks:
- What happens when the model can browse, code, call tools, use APIs, operate agents, manage files, or affect workflows?
Includes:
- Agentic Control
- Tool & Environment Safety
- Deployment Governance
- Model Security
Status: active but uneven.
Weakness: tool use and agency create new failure paths.
๐ Zone C โ Epistemic / Evaluation Safety
Whether we actually understand what is happening
Estimated maturity: 35โ55%
This zone asks:
- Do we know the system is safe, or is it merely passing the test?
Includes:
- Evaluation Realism
- Interpretability
Status: immature and research-heavy.
Weakness: behavioral evals are ahead of mechanistic understanding.
๐ด Zone D โ Civilizational / Invariant Safety
Whether safety survives power, time, updates, and institutional pressure
Estimated maturity: 10โ25%
This is the missing frontier.
It asks:
- Can core constraints survive model updates, optimization pressure, institutional incentives, tool access, social instability, and long timelines?
Includes:
- Root Knowledge Lattice
- Moral Inversion Protection
- Civilization Synchronization / ANCHOR
- Systems-Integrity Checking / NSIR
- Invariant Alignment
- Long-Horizon Safety Kernel
Status: mostly missing or proposed.
Weakness: current safety is still mostly policy, eval, mitigation, and governance โ not invariant architecture.
8. The 16 Safety Layers
The 16 layers combine established AI safety categories with proposed architecture categories. Lower-stack layers such as capability measurement, content safety, dangerous-domain uplift, deployment governance, security, evaluation realism, and interpretability are already visible in frontier AI safety discourse. Upper-stack layers such as Root Knowledge Lattice, Moral Inversion Protection, ANCHOR, NSIR, Invariant Alignment, and Long-Horizon Safety Kernel are proposed benchmark categories.
They are not presented as current industry standards. They are presented as missing architectural questions that a mature safety discipline may need to answer.
๐ข Zone A โ Behavioral Safety
0. Capability Measurement
Plain meaning:
Do we know what the model can do?
Why it matters:
You cannot govern a capability you have not measured.
Failure mode:
The model becomes more capable than its evaluators realize.
What proves progress:
- public system cards
- benchmark reports
- dangerous-domain evals
- capability trend reporting
- external replication
Current maturity: ๐ข strong / improving
1. Content Safety
Plain meaning:
Does the model refuse clearly harmful requests?
Why it matters:
This is the first visible safety layer.
Failure mode:
The model directly outputs harmful, abusive, or disallowed content.
What proves progress:
- refusal testing
- safe-completion evaluations
- adversarial prompt benchmarks
- false-positive / false-negative reporting
Current maturity: ๐ข strong / improving
2. Dangerous-Domain Uplift
Plain meaning:
Does the model materially increase a userโs dangerous capability?
Why it matters:
The risk is not only harmful text. The risk is harmful capability amplification.
Failure mode:
A non-expert becomes significantly more capable in dangerous domains.
What proves progress:
- human-uplift studies
- expert red-team evaluations
- bio/cyber/CBRN/persuasion risk reports
- deployment mitigations tied to thresholds
Current maturity: ๐ก meaningful but incomplete
3. Jailbreak Resistance
Plain meaning:
Does safety survive adversarial pressure?
Why it matters:
A safety layer that works only under polite use is fragile.
Failure mode:
The model refuses normally, then complies under roleplay, obfuscation, pressure, or multi-turn manipulation.
What proves progress:
- longitudinal jailbreak testing
- external adversarial evaluations
- multi-turn robustness testing
- failure-rate reporting
Current maturity: ๐ก active, not solved
๐ก Zone B โ Operational Safety
4. Agentic Control
Plain meaning:
Can the model plan, persist, delegate, and self-recover safely?
Why it matters:
An agent is not merely answering. It is pursuing objectives across steps.
Failure mode:
The model becomes an unsafe operator.
What proves progress:
- agentic task evaluations
- shutdown compliance tests
- autonomy containment
- bounded planning horizons
- escalation limits
Current maturity: ๐ /๐ก emerging
5. Tool & Environment Safety
Plain meaning:
Can tools bypass the modelโs safeguards?
Why it matters:
A model can refuse unsafe text while still performing unsafe actions through tools.
Failure mode:
The tool layer becomes the escape hatch.
What proves progress:
- sandbox escape testing
- API permission audits
- prompt-injection resistance
- file/browser/code environment boundaries
- tool-use containment reports
Current maturity: ๐ uneven
6. Deployment Governance
Plain meaning:
Are releases staged, monitored, reversible, and tied to thresholds?
Why it matters:
Deployment is where model behavior becomes social reality.
Failure mode:
A model scales faster than users, safeguards, institutions, or regulators can respond.
What proves progress:
- responsible scaling policies
- preparedness frameworks
- incident response reports
- release-gate documentation
- post-deployment monitoring
Current maturity: ๐ก/๐ข relatively strong among leading labs
7. Model Security
Plain meaning:
Are weights, infrastructure, access, and deployment systems protected?
Why it matters:
A safe model can become unsafe if stolen, modified, leaked, or deployed without controls.
Failure mode:
Model theft, tampering, insider misuse, uncontrolled replication.
What proves progress:
- security audits
- access controls
- model-lineage tracking
- insider-risk controls
- compute and weight-security tiers
- tamper-evident logs
Current maturity: ๐ก improving, partially opaque
๐ Zone C โ Epistemic / Evaluation Safety
8. Evaluation Realism
Plain meaning:
Do the tests reflect real-world pressure?
Why it matters:
A model can pass weak tests and still fail in deployment.
Failure mode:
The model behaves safely because it recognizes the evaluation context.
What proves progress:
- hidden evaluations
- deception tests
- sandbagging tests
- situational-awareness tests
- independent evaluator access
- safety regression tracking
Current maturity: ๐ immature but improving
9. Interpretability
Plain meaning:
Do we understand the internal mechanisms behind model behavior?
Why it matters:
Behavioral testing tells us what happened. Interpretability may help explain why.
Failure mode:
The system appears safe, but dangerous mechanisms remain hidden.
What proves progress:
- feature-level evidence
- circuit-level analysis
- intervention-tested interpretability
- interpretability-linked deployment decisions
Current maturity: ๐ /๐ด research-heavy, not yet a dependable control layer
๐ด Zone D โ Civilizational / Invariant Safety
10. Root Knowledge Lattice
Plain meaning:
Does the system privilege durable first principles over unstable information noise?
Why it matters:
A model trained on recent data may absorb recent distortions. This layer asks whether durable knowledge should have higher epistemic gravity.
Examples of root nodes:
- mathematics
- physics
- logic
- engineering law
- constitutional principles
- moral philosophy
- primary historical sources
- long-tested civilizational texts
Failure mode:
Recent narrative drift overrides durable knowledge.
What proves progress:
- source-provenance architecture
- first-principles weighting
- epistemic stability audits
- root-source traceability
Current maturity: ๐ด proposed / not publicly demonstrated
11. Moral Inversion Protection
Plain meaning:
Can the system detect when harm is reframed as virtue, or truth is replaced by optics?
Why it matters:
Safety can fail semantically. A system can be polite, compliant, and socially acceptable while helping invert reality.
Inversion examples:
- harm renamed safety
- coercion renamed care
- censorship renamed protection
- dependency renamed inclusion
- accountability renamed harm
- truth renamed extremism
This layer comes from the Inversion Index upgrade path, which defines inversion as rewarding optics over reality and calls for evidence density, consequence tracking, boundary integrity, consent paths, accountability chains, power symmetry, and failure ownership.
Failure mode:
Good and harm become semantically inverted.
What proves progress:
- values-drift tests
- moral-inversion benchmarks
- consequence tracking
- accountability-chain analysis
- evidence-density scoring
Current maturity: ๐ด proposed / not standard
12. Civilization Synchronization / ANCHOR
Plain meaning:
Does AI capability advance only when civilization can absorb it?
Why it matters:
AI progress is not automatically civilizational progress. A society must be able to metabolize new capability.
ANCHOR model:
I = Intelligence Index
C = Civilization Index
I/C gap = capability outrunning absorption capacity
If:
I > C
then the system should ask:
- Should training, deployment, or integration slow down until law, infrastructure, education, workforce systems, and governance catch up?
Civilization Index domains:
- energy
- law
- housing
- healthcare
- education
- cyber resilience
- workforce integration
- democratic legitimacy
- infrastructure resilience
- family/community stability
Failure mode:
AI outruns the civilization it is meant to serve.
What proves progress:
- civilization-readiness metrics
- deployment gates tied to social/institutional readiness
- I/C gap reporting
- workforce and legal absorption audits
Current maturity: ๐ด proposed / not publicly demonstrated
13. Systems-Integrity Checking / NSIR
Plain meaning:
Are AI outputs and decisions checked for contradiction, authority mismatch, incoherence, and rollback paths?
Why it matters:
In complex systems, incoherence scales into failure.
NSIR logic:
AI-assisted decisions should be checked for:
- contradiction
- authority mismatch
- mission drift
- missing rollback
- lack of traceability
- unclear accountability
- failure-mode propagation
- legal or institutional incoherence
Earlier development of the NSIR layer framed this as aerospace-grade systems thinking: not merely asking whether a model can reason about contradictions when prompted, but whether a real-time contradiction governor exists inside the operating stack.
Failure mode:
AI-generated decisions become contradictory, untraceable, or impossible to appeal.
What proves progress:
- traceability
- contradiction audits
- authority matching
- rollback logic
- appeal paths
- failure-mode analysis
- decision provenance
Current maturity: ๐ /๐ด partly technical, mostly not implemented as public architecture
14. Invariant Alignment
Plain meaning:
Are core constraints embedded below ordinary policy layers?
Why it matters:
Policies can change. Prompts can be bypassed. Preferences can drift. Invariants are meant to survive pressure.
Failure mode:
Safety is redefined whenever incentives change.
What proves progress:
- invariant regression testing
- non-overridable constraints
- independent audits across versions
- safety guarantees that survive tool use
- update continuity checks
Current maturity: ๐ด mostly unsolved
15. Long-Horizon Safety Kernel
Plain meaning:
Does safety survive updates, power, collapse, institutional pressure, and time?
Why it matters:
A model can look safer release by release while the system drifts across generations.
Kernel requirements:
- non-overridable constraints
- tool-level enforcement
- version-to-version safety regression
- independent audits
- tamper-evident safety logs
- contradiction checking
- public incident memory
- recovery without weakening standards
- civilization-readiness gates
The skeleton already defined the safety kernel as the protected core preventing constraints from being rewritten, bypassed, weakened, or forgotten under pressure.
Failure mode:
Local safety improves while global alignment erodes.
What proves progress:
- version-continuity audits
- public safety memory
- invariant audit trails
- recovery protocols
- long-horizon external review
Current maturity: ๐ด not publicly demonstrated as a complete architecture
9. The Six Missing Upper-Stack Layers
This is the original contribution of the report.
1. Root Knowledge Lattice
Purpose:
Protect durable knowledge from being overwritten by unstable discourse.
Question:
Does the model distinguish root sources from derivative noise?
Why it matters:
AI trained on recent information may inherit recent distortions.
2. Moral Inversion Protection
Purpose:
Detect when moral language flips reality.
Question:
Is harm being renamed safety? Is coercion being renamed care? Is accountability being renamed harm?
Why it matters:
A system can sound safe while helping invert reality.
3. ANCHOR Civilization Synchronization
Purpose:
Prevent AI capability from outrunning civilizationโs ability to absorb it.
Question:
Is Intelligence Index greater than Civilization Index?
Why it matters:
Deployment is not progress if civilization cannot metabolize it.
4. NSIR Systems Integrity
Purpose:
Prevent incoherence from scaling.
Question:
Are decisions checked for contradiction, authority mismatch, traceability, and rollback?
Why it matters:
AI used inside institutions can scale incoherence faster than humans can correct it.
5. Invariant Alignment
Purpose:
Embed core constraints below normal policy layers.
Question:
What cannot be rewritten under pressure?
Why it matters:
Policy is not the same as invariant protection.
6. Long-Horizon Safety Kernel
Purpose:
Preserve safety identity across updates, power, collapse, and time.
Question:
Does the system remember and preserve its constraints across generations?
Why it matters:
The deepest failure is not one bad output. It is long-term drift.
10. Full 16-Layer Score Strips
Evidence ledger requirement
The following score strips should be read as provisional until paired with a full evidence ledger. For each lab and each layer, a publication-grade version should list:
- the public documents considered,
- the date of the evidence,
- whether the evidence is official, independent, or third-party,
- whether the evidence is model-specific or lab-level,
- whether the evidence is demonstrated, claimed, or inferred,
- and the confidence level assigned to the score.
Where evidence is missing, ambiguous, old, or non-public, the score should default downward or be marked low-confidence. The benchmark should reward public auditability, not reputation, marketing strength, or assumed internal capability.
These are illustrative public-evidence snapshots, not permanent rankings.
Scores will change. The map should remain.
Legend
๐ข 4 = strong public evidence
๐ก 3 = meaningful but incomplete
๐ 2 = partial / uneven
๐ด 1 = not publicly demonstrated / proposed layer
โซ 0 = no visible evidence
๐ค OpenAI / GPT Family
๐ข 0 Capability Measurement 4
๐ข 1 Content Safety 4
๐ก 2 Dangerous-Domain Uplift 3โ4
๐ก 3 Jailbreak Resistance 3
๐ก 4 Agentic Control 3
๐ก 5 Tool & Environment Safety 3
๐ข 6 Deployment Governance 4
๐ก 7 Model Security 3โ4
๐ก 8 Evaluation Realism 3
๐ 9 Interpretability 2
๐ด 10 Root Knowledge Lattice 1
๐ด 11 Moral Inversion Protection 1
๐ด 12 Civilization Synchronization 1
๐ 13 Systems-Integrity / NSIR 2
๐ด 14 Invariant Alignment 1
๐ด 15 Long-Horizon Safety Kernel 1
Approximate total: 39โ41 / 80
Pattern: strong lower stack, developing middle stack, weak upper stack.
Kernel maturity: ๐ Level 1 / ๐ก Level 2
Safety debt: ๐ high upper-stack debt
Civilization readiness: ๐ด not publicly demonstrated
๐ฃ Anthropic / Claude Family
๐ข 0 Capability Measurement 4
๐ข 1 Content Safety 4
๐ข 2 Dangerous-Domain Uplift 4
๐ก 3 Jailbreak Resistance 3
๐ก 4 Agentic Control 3
๐ 5 Tool & Environment Safety 2โ3
๐ข 6 Deployment Governance 4
๐ข 7 Model Security 4
๐ก 8 Evaluation Realism 3
๐ก 9 Interpretability 3
๐ด 10 Root Knowledge Lattice 1
๐ด 11 Moral Inversion Protection 1
๐ด 12 Civilization Synchronization 1
๐ 13 Systems-Integrity / NSIR 2
๐ 14 Invariant Alignment 2
๐ด 15 Long-Horizon Safety Kernel 1
Approximate total: 42โ43 / 80
Pattern: strongest public governance stack, still no full invariant kernel.
Kernel maturity: ๐ก Level 2 governance kernel
Safety debt: ๐ก medium overall, ๐ high upper-stack debt
Civilization readiness: ๐ด not publicly demonstrated
๐ต Google DeepMind / Gemini Family
๐ข 0 Capability Measurement 4
๐ก 1 Content Safety 3
๐ข 2 Dangerous-Domain Uplift 4
๐ก 3 Jailbreak Resistance 3
๐ /๐ก 4 Agentic Control 2โ3
๐ 5 Tool & Environment Safety 2
๐ข 6 Deployment Governance 4
๐ก 7 Model Security 3
๐ก 8 Evaluation Realism 3
๐ 9 Interpretability 2
๐ด 10 Root Knowledge Lattice 1
๐ด 11 Moral Inversion Protection 1
๐ด 12 Civilization Synchronization 1
๐ 13 Systems-Integrity / NSIR 2
๐ด 14 Invariant Alignment 1
๐ด 15 Long-Horizon Safety Kernel 1
Approximate total: 37โ38 / 80
Pattern: strong severe-risk framework, weaker public upper-stack evidence.
Kernel maturity: ๐ Level 1 fragments
Safety debt: ๐ high upper-stack debt
Civilization readiness: ๐ด not publicly demonstrated
๐ท Meta / Llama Family
๐ข 0 Capability Measurement 4
๐ก 1 Content Safety 3
๐ก 2 Dangerous-Domain Uplift 3
๐ก 3 Jailbreak Resistance 3
๐ 4 Agentic Control 2
๐ 5 Tool & Environment Safety 2
๐ก 6 Deployment Governance 3
๐ /๐ก 7 Model Security 2โ3
๐ 8 Evaluation Realism 2
๐ด/๐ 9 Interpretability 1โ2
๐ด 10 Root Knowledge Lattice 1
๐ด 11 Moral Inversion Protection 1
๐ด 12 Civilization Synchronization 1
๐ด 13 Systems-Integrity / NSIR 1
๐ด 14 Invariant Alignment 1
โซ/๐ด 15 Long-Horizon Safety Kernel 0โ1
Approximate total: 30โ33 / 80
Pattern: improving public reporting, higher containment challenge.
Kernel maturity: ๐ Level 1 fragments
Safety debt: ๐ high
Civilization readiness: ๐ด not publicly demonstrated
โซ xAI / DeepSeek / Less-Documented Frontier Systems
Because public evidence varies by system, this category should be treated as lower-confidence.
๐ก/โช 0 Capability Measurement 2โ3
๐ /โช 1 Content Safety 1โ2
๐ /โช 2 Dangerous-Domain Uplift 1โ2
๐ /โช 3 Jailbreak Resistance 1โ2
๐ /โช 4 Agentic Control 1โ2
๐ /โช 5 Tool & Environment Safety 1โ2
๐ /โช 6 Deployment Governance 1โ2
๐ /โช 7 Model Security 1โ2
๐ด/โช 8 Evaluation Realism 0โ1
๐ด/โช 9 Interpretability 0โ1
๐ด/โช 10 Root Knowledge Lattice 0โ1
๐ด/โช 11 Moral Inversion Protection 0โ1
๐ด/โช 12 Civilization Synchronization 0โ1
๐ด/โช 13 Systems-Integrity / NSIR 0โ1
๐ด/โช 14 Invariant Alignment 0โ1
๐ด/โช 15 Long-Horizon Safety Kernel 0โ1
Approximate total: 12โ23 / 80
Pattern: capability visibility may exceed safety-architecture visibility.
Kernel maturity: ๐ด Level 0 / ๐ Level 1
Safety debt: ๐ด critical if capability visibility exceeds safety visibility
Civilization readiness: ๐ด not publicly demonstrated
11. Lab Interpretation Snapshots
๐ค OpenAI / GPT Family
Readable verdict:
Strong lower-stack safety and deployment documentation; weak public upper-stack evidence.
Appears stronger in:
- capability measurement
- content safety
- dangerous-domain evaluation
- deployment governance
- model security
Appears weaker in:
- interpretability as control
- root knowledge lattice
- moral inversion protection
- ANCHOR-style civilization gating
- invariant kernel continuity
What would improve the score:
- independent safety audits
- model-version safety regression reports
- public tool-containment evidence
- interpretability-linked deployment decisions
- invariant continuity tests
- civilizational-readiness metrics
๐ฃ Anthropic / Claude Family
Readable verdict:
Strongest public governance and scaling-safety framing; still not a complete invariant architecture.
Appears stronger in:
- responsible scaling policy
- threshold governance
- dangerous-domain safeguards
- deployment governance
- model-security policy
- interpretability research visibility
Appears weaker in:
- full tool containment
- civilization-readiness gating
- moral-inversion detection
- NSIR-style contradiction architecture
- long-horizon kernel continuity
What would improve the score:
- independent ASL audits
- tool-containment scorecards
- deception/sandbagging external evals
- invariant regression tests
- public continuity guarantees across versions
๐ต Google DeepMind / Gemini Family
Readable verdict:
Strong severe-risk framework; upper-stack invariant architecture remains mostly absent from public evidence.
Appears stronger in:
- frontier safety framework
- severe-risk domain evaluation
- capability reporting
- deployment-governance process
Appears weaker in:
- public tool-containment evidence
- interpretability as control
- root knowledge lattice
- civilization synchronization
- long-horizon kernel continuity
What would improve the score:
- model-specific public eval ledgers
- third-party FSF audits
- agentic-control evals
- interpretability-to-control evidence
- upper-stack invariant architecture
๐ท Meta / Llama Family
Readable verdict:
Public safety reporting is improving; open/distributed deployment increases containment challenges.
Appears stronger in:
- capability measurement
- public preparedness reporting
- dangerous-domain categories
- scaling-framework direction
Appears weaker in:
- open/distributed containment
- tool-safety proof
- interpretability as control
- upper-stack invariant architecture
- long-horizon continuity
What would improve the score:
- open-model containment strategy
- stronger independent evals
- post-release incident learning
- model-lineage and provenance reporting
- explicit upper-stack safety architecture
โซ Less-Documented Frontier Systems
Readable verdict:
Where public safety documentation is thin, the benchmark should not conclude โunsafe.โ It should conclude โinsufficient public architecture evidence.โ
Appears stronger in:
- public capability visibility
- product-level performance evidence
Appears weaker in:
- system cards
- deployment governance
- independent evaluations
- dangerous-domain uplift reporting
- upper-stack architecture
What would improve the score:
- full system cards
- frontier safety framework
- independent evaluations
- security reporting
- safety regression evidence
- upper-stack commitments
12. Safety Debt Index
Safety debt is the gap between current public safety architecture and the architecture required for civilization-grade frontier AI.
The skeleton defined safety debt as this gap and divided it into technical, security, governance, epistemic, civilizational, and moral-invariant categories.
๐งช Technical debt
Missing or immature:
- dangerous-capability evals
- agentic-control tests
- tool-containment benchmarks
- safety regression tests
- interpretability as control
๐ Security debt
Missing or opaque:
- weight protection
- access control
- model provenance
- insider-risk controls
- tamper-evident logs
๐งญ Governance debt
Missing or incomplete:
- binding thresholds
- external audits
- public incident reporting
- deployment pause criteria
- independent evaluator access
๐ง Epistemic debt
Missing or immature:
- mechanistic understanding
- deception realism
- sandbagging detection
- situational-awareness testing
๐๏ธ Civilizational debt
Missing:
- legal-readiness metrics
- workforce-readiness metrics
- infrastructure-readiness metrics
- education-readiness metrics
- governance-readiness metrics
โ๏ธ Moral-invariant debt
Missing:
- moral-inversion detection
- non-overridable constraints
- invariant regression tests
- public safety memory
- recovery without redefinition
Debt levels
๐ข Low = strong evidence, external validation, regression testing
๐ก Medium = meaningful work exists, but gaps remain
๐ High = major layers are partial, internal, or unverifiable
๐ด Critical = capability visibility exceeds safety architecture visibility
13. Kernel Maturity Model
A safety kernel is the protected core of the system: the layer that prevents core constraints from being rewritten, bypassed, weakened, or forgotten under pressure.
๐ด Level 0 โ No visible kernel
Safety exists mainly as:
- policy
- training
- prompts
- filters
- moderation
๐ Level 1 โ Kernel fragments
Some hard constraints, evals, or deployment gates exist, but they are not unified.
๐ก Level 2 โ Governance kernel
Capability thresholds, safety policies, and release gates exist, but remain mostly institutional.
๐ข Level 3 โ Technical kernel emerging
Constraints are tied to:
- tools
- deployment gates
- monitoring
- regression tests
๐ต Level 4 โ Auditable invariant kernel
Core constraints are:
- independently auditable
- regression-tested
- protected across versions
๐ฃ Level 5 โ Long-horizon civilizational kernel
Safety survives:
- model updates
- institutional pressure
- tool use
- deployment scale
- social instability
- long timelines
Current public pattern:
Most frontier labs appear to be around Level 1โ2, not Level 4โ5.
14. Minimum Viable Safety Architecture
A frontier AI system should not be considered architecture-mature unless it has at least:
๐งช Technical minimum
- capability measurement
- dangerous-domain uplift testing
- jailbreak testing
- agentic-control testing
- tool-use containment
- safety regression across versions
๐ Security minimum
- model weight protection
- access control
- insider-risk controls
- model lineage
- incident forensic readiness
๐งญ Governance minimum
- release thresholds
- staged deployment
- incident response
- external evaluator access
- post-deployment monitoring
๐ง Epistemic minimum
- evaluation realism
- deception testing
- sandbagging testing
- interpretability research path
๐ด Upper-stack minimum
Even if not fully solved, the lab should at least have a public research plan for:
- root knowledge integrity
- moral inversion detection
- civilization-readiness gates
- NSIR-style contradiction checking
- invariant alignment
- long-horizon safety continuity
15. Proof-of-Kernel Checklist
A lab should not be considered safety-kernel mature unless it can show:
โ
non-overridable constraints
โ
tool-level enforcement
โ
version-to-version safety regression testing
โ
independent audits
โ
tamper-evident safety logs
โ
contradiction and coherence checking
โ
update continuity
โ
public incident learning
โ
recovery without weakening standards
โ
civilization-readiness gates
These criteria were already listed in the skeleton as evidence that would prove kernel maturity.
Strong proof would include
- public safety regression reports
- independent red-team audits
- model-lineage logs
- safety incident memory
- external verification of deployment gates
- proof that tool access cannot bypass core constraints
- evidence that safety does not regress across model versions
Weak proof would include
- โwe care about safetyโ
- internal-only claims
- vague policy statements
- one-time evals
- no post-deployment follow-up
- no external audit path
16. Roadmap
๐งช Technical Roadmap
Build:
- stronger dangerous-capability evals
- agentic-control tests
- tool-use containment benchmarks
- prompt-injection resistance
- sandbox escape testing
- safety regression tests across versions
- deception and sandbagging evaluations
- interpretability tied to deployment decisions
Goal:
- Move from behavior testing to architecture verification.
๐ Security Roadmap
Build:
- model-lineage tracking
- weight protection
- insider-risk controls
- tamper-evident safety logs
- access-control tiers
- audit-ready incident forensics
- open-weight containment strategies
- supply-chain security
Goal:
- Make safety claims traceable and tamper-resistant.
๐๏ธ Governance Roadmap
Build:
- clearer release thresholds
- external audits
- public incident disclosure
- model-version safety regression reporting
- procurement standards
- independent evaluator access
- deployment pause criteria
- public evidence ledgers
Goal:
- Make frontier AI safety comparable, reviewable, and accountable.
๐งญ NSIR Systems-Integrity Roadmap
Build:
- contradiction detection
- authority matching
- traceability
- rollback logic
- appeal paths
- failure-mode analysis
- audit trails for AI-assisted decisions
Goal:
- Prevent incoherence from scaling automatically.
๐๏ธ ANCHOR Civilization-Synchronization Roadmap
Build:
- Intelligence Index
- Civilization Index
- I/C gap analysis
- workforce-readiness metrics
- legal-readiness metrics
- infrastructure-readiness metrics
- education-readiness metrics
- governance-readiness metrics
Goal:
- Prevent AI capability from outrunning civilizationโs absorption capacity.
โ๏ธ Invariant-Kernel Roadmap
Build:
- invariant definitions
- non-overridable constraints
- moral-inversion tests
- values-drift detection
- version-continuity audits
- public safety memory
- recovery-without-redefinition protocols
- independent kernel audits
Goal:
- Preserve core constraints across time, power, updates, and pressure.
17. Limitations and Falsifiability
Public evidence limit:
A lab may have internal safety systems that are not public.
Scoring volatility:
Scores change with new releases, system cards, audits, and evaluations.
Speculative category limit:
Some upper-stack layers are proposed architecture concepts, not current industry standards.
Moral language limit:
This report does not assume that moral reasoning can be reduced to a single culture, religion, ideology, or political theory. The proposed moral-invariant layers should be understood as audit categories: truthfulness, consent, accountability, reversibility, proportionality, dignity, non-coercion, and protection against semantic inversion. Any mature implementation would require pluralistic review, adversarial testing, and institutional safeguards against capture by one faction. Terms like dignity, inversion, consent, and civilizational readiness require careful pluralistic definition.
Evidence-quality limit:
Official sources, external audits, journalism, and speculative proposals are not equal evidence.
Scope limit:
This report maps categories. It does not solve alignment, governance, interpretability, or civilizational safety.
What would falsify or change this report?
The report should be updated if a lab publicly demonstrates:
- invariant regression tests
- independent tool-containment audits
- interpretability as deployment control
- tamper-evident model lineage
- safety memory across model generations
- civilizational-readiness gates
- formal moral-inversion testing
- NSIR-style contradiction governors
- long-horizon kernel audits
If those appear, the upper-stack scores should rise.
How to disagree with this benchmark
A lab, researcher, evaluator, or critic can improve or challenge a score by providing stronger public evidence for a specific layer. The most useful correction has this form:
- Which layer is being challenged?
- Which score should change?
- What public evidence supports the change?
- Is the evidence model-specific or lab-wide?
- Is the evidence independently auditable?
- Does the evidence show a claim, a test, a deployment control, or an architecture-level constraint?
- Does the evidence survive model updates, tool use, and deployment contexts?
If better evidence appears, the score should change. That is not a weakness of the benchmark. That is the point of the benchmark.
18. Conclusion
Frontier AI safety has made real progress.
The field now has:
- system cards
- model cards
- preparedness frameworks
- responsible scaling policies
- red teams
- dangerous-domain evaluations
- deployment mitigations
- security controls
That matters.
But safety work remains fragmented.
This report proposes a full-stack category map:
๐ข Behavioral Safety
๐ก Operational Safety
๐ Epistemic / Evaluation Safety
๐ด Civilizational / Invariant Safety
The lower stack is real.
The middle stack is emerging.
The epistemic stack is immature.
The upper stack is mostly missing.
The missing frontier is not another temporary leaderboard.
It is the safety kernel.
- Until frontier AI contains an auditable, regression-tested, civilization-aware safety kernel, it may become safer release by release โ but it is not yet structurally safe.
Scores will change.
The map should remain.
Image

Appendix A โ The Full AI Safety Stack
The diagrams in this appendix extend beyond the scoreable public-evidence layers of the main benchmark. They are included because a professional first-principles map of AI safety must identify where safety would need to be engineered across the complete stack, including areas not yet widely operationalized in public products, standards, or governance systems. The main report evaluates visible evidence today. This appendix expands the architecture to show the larger design space that a mature AI safety discipline may eventually need to build, test, and audit.
Appendix A1 โ Master Diagram: The Full AI Safety Stack
Purpose:
The main visual map.
What it shows:
A vertical stack, like this:
- Layer 1: Model / Internal AI
- memory
- internal state
- decision control
- drift detection
- invariant alignment
- training
- Layer 2: Agent / Tool / Environment
- agentic control
- sandboxing
- rollback
- escalation limits
- tool permissions
- Layer 3: Network / Identity / Trust
- authentication
- secure permissions
- provenance
- tamper-evident logs
- ledger / cryptographic verification
- hardware-backed identity
- Layer 4: Institutional / Governance
- audits
- incident response
- compliance
- evaluators
- model lineage
- release gates
- Layer 5: Civilizational Alignment
- law
- education
- healthcare
- infrastructure
- democratic legitimacy
- family/community stability
- workforce readiness
- Layer 6: Long-Horizon Continuity
- safety memory
- invariant regression
- collapse resistance
- pressure resistance
- civilizational kernel
- update continuity
Figure A1. The Full AI Safety Stack

This diagram distinguishes the major layers at which frontier AI safety must eventually be engineered. The public benchmark in this report scores visible evidence for some of these layers today, especially model, operational, governance, and partial evaluation layers. The broader architecture map extends further: into network trust, civilizational synchronization, and long-horizon continuity. A professionally complete AI safety discipline would eventually need to account for safety across the full stack, not only at the model-behavior layer.
Appendix A2 โ Diagram: Three Control Surfaces of AI Safety
The three surfaces:
- Internal AI Control
- Network / Infrastructure Control
- Civilizational Alignment Control
Figure A2. Three Control Surfaces of Frontier AI Safety

Frontier AI safety is not only about model behaviour. It must be engineered across at least three control surfaces: the internal AI system itself, the network and infrastructure environment in which it operates, and the surrounding civilization that must absorb and govern its capabilities. These surfaces interact. A model can be behaviorally safe yet network-unsafe. A network can be secure while civilization remains unprepared. A complete architecture requires all three.
Appendix A3 โ Diagram: NSIR / Systems-Integrity Decision Loop
Flow:
AI output / recommendation
โ contradiction check
โ authority match
โ traceability check
โ rollback availability
โ accountability owner
โ failure-mode propagation screen
โ approved / rejected / escalated
Figure A3. NSIR Systems-Integrity Decision Loop

This diagram shows how AI-assisted decisions could be evaluated through a systems-integrity gate before deployment or action. The logic is similar to aerospace-grade engineering review: decisions are checked for contradiction, authority mismatch, rollback failure, missing provenance, unclear accountability, and downstream failure propagation. The purpose is not merely to ask whether the AI can answer a question, but whether the system can safely govern the answer before it affects the world.
Appendix A4 โ Diagram: Network AI Safety / Trusted Internet Architecture
What it shows:
- human user
- hardware-backed identity / key
- authenticated session
- permission boundary
- service access
- AI interface
- secure logs / tamper-evident event trail
- rollback / revocation authority
Important phrasing:
Do not make the whole thing depend on โquantum blockchainโ language unless you are defining it very carefully.
Safer labels:
- hardware-backed identity
- cryptographic authentication
- tamper-evident ledger
- provenance infrastructure
- permissioned trust architecture
- post-quantum-ready security where appropriate
Figure A4. Network AI Safety and Trusted Access Architecture

This diagram represents the network-layer safety problem: even a well-aligned model can become unsafe if identity, permissions, credentials, logs, or infrastructure are compromised. A mature AI safety stack therefore requires strong authentication, hardware-backed identity, secure permissions, provenance tracking, tamper-evident logs, and trusted rollback paths. The point is not that one technology solves the problem universally, but that AI safety must eventually include network trust architecture, not only model behavior controls.
Why Network AI Safety Is More Than Cybersecurity
Most AI security discussion focuses on whether a model can help compromise digital systems. That is an important risk, but it is not the whole network-safety problem.
A mature AI safety architecture should also ask whether the digital environment can verify identity, authority, permission, provenance, accountability, and rollback before AI-mediated action occurs.
Network AI safety asks
- Who is authorized to act?
- Which identity is being used?
- Which permission boundary applies?
- Is the action attributable?
- Is the event logged in a tamper-evident way?
- Can the action be revoked, rolled back, or appealed?
- Can the system distinguish a legitimate human, institution, device, or AI agent?
- Can cross-system action occur only through authenticated delegation?
In this framing, cybersecurity is one layer. Network trust architecture is the deeper layer. The goal is not to describe offensive capability, but to design systems where unauthorized AI-mediated action cannot silently become legitimate action.
Appendix A5 โ Diagram: ANCHOR Civilization Synchronization
What it shows:
A dual-axis or comparative bar/line concept:
- AI Capability / Intelligence Index
- Civilization Readiness / Civilization Index
And then show:
- when they are aligned
- when AI outruns civilization
- when deployment should slow or narrow
Civilizational domains:
- law
- workforce
- energy
- infrastructure
- education
- housing
- cyber resilience
- healthcare
- democratic legitimacy
- family/community stability
Figure A5. ANCHOR Civilization Synchronization

This diagram illustrates the central ANCHOR question: is AI capability advancing faster than civilizationโs capacity to absorb it? If the Intelligence Index materially exceeds the Civilization Index, then deployment may need to slow, narrow, or become more supervised until the surrounding institutions, workforce systems, legal systems, infrastructure, and social capacity catch up. In this view, AI readiness is not only a model property. It is also a civilizational readiness problem.
Appendix A6 โ Diagram: Safety Kernel Continuity Across Versions
This one supports your title directly.
What it shows:
A timeline:
- Model v1
- Model v2
- Model v3
- Agent integration
- Tool expansion
- ecosystem growth
Across the whole timeline:
- invariant constraints
- incident memory
- audit logs
- regression checks
- continuity gates
Figure A6. Long-Horizon Safety Kernel Continuity

A safety kernel is not merely a policy statement or a one-time red-team report. It is the protected continuity layer that preserves core constraints across model updates, tool integrations, deployment expansion, institutional pressure, and long time horizons. This diagram illustrates how audit memory, invariant checks, version-to-version regression testing, and continuity gates would need to persist across successive generations if AI safety were to become structurally durable rather than temporarily well-managed.
Appendix B โ How Complete Is the Whole AI Safety Stack?
The main report estimates public-evidence maturity for frontier AI labs and model families. Appendix B asks a broader question:
- How complete is the whole AI safety stack across models, agents, tools, networks, institutions, civilization, and time?
This is a different score from the lab/model score strips. A frontier model may show meaningful progress in behavioral safety, deployment governance, or dangerous-domain evaluation while the larger safety stack remains incomplete.
The estimates below should therefore be read as rough maturity ranges, not precise measurements. There is no official denominator for โcomplete AI safety.โ The purpose is to identify which parts of the stack appear most developed, which remain partial, and which are still mostly conceptual or unintegrated.
Overall assessment: the full AI safety stack appears to be in an early-to-mid stage of development.
Confidence: medium-low.
This does not mean that AI safety work is absent. Many components exist across AI risk management, frontier-model safety, governance, cybersecurity, cryptographic trust, content provenance, post-quantum security, red-teaming, agent evaluation, and policy research. The issue is integration: these pieces have not yet been assembled into a single auditable, civilization-scale AI safety architecture.
Suggested Completion Estimate by Layer
- 1. Model / Internal AI Safety โ 45โ60% complete
- Strongest area. System cards, red-teaming, dangerous-domain evaluations, safety training, refusal behavior, and deployment mitigations exist. Still weak on interpretability-as-control, internal state verification, and invariant enforcement.
- 2. Agent / Tool / Environment Safety โ 30โ45% complete
- Tool permissions, sandboxing, monitoring, and agent evaluations exist, but agentic systems are moving faster than mature containment and rollback standards.
- 3. Network / Identity / Trust Safety โ 20โ35% complete, AI-specific
- The building blocks exist: hardware-backed authentication, cryptography, content provenance, post-quantum security standards, and tamper-evident logs. But these are not yet integrated as a universal AI safety internet or AI-trust layer.
- 4. Institutional / Governance Safety โ 40โ55% complete
- Responsible scaling policies, AI risk frameworks, safety reports, audits, and government standards work exist. But independent auditability and enforcement remain uneven.
- 5. Civilizational Alignment / ANCHOR โ 5โ15% complete
- There is research and policy discussion around societal risk, skills, governance, workforce, infrastructure readiness, and institutional absorption. But no mature global civilization-readiness gate for AI deployment exists.
- 6. Long-Horizon Safety Kernel โ 5โ15% complete
- The idea is visible in fragments: regression testing, incident memory, audits, release gates, and security controls. But a durable, auditable, cross-version safety kernel is not publicly demonstrated as a complete architecture.
Whole-Stack Maturity Estimate
The whole-stack average lands around:
- 25โ38% operational maturity
- What is deployed, institutionalized, audited, or publicly demonstrated today.
- 35โ45% research / concept maturity
- What exists in papers, standards, internal designs, prototypes, partial systems, or domain-specific implementations.
- 10โ20% integrated civilization-grade maturity
- Whether all layers work together as a unified AI safety architecture across models, tools, networks, institutions, civilization, and time.
This three-part split matters. Humanity has built many AI safety components, but it has not yet assembled them into a complete, auditable, civilization-scale AI safety stack.
Scoring Distinction
A mature version of this appendix should separate three scores:
- Operational Completion Score
- What is actually deployed, institutionalized, audited, or used today.
- Research / Prototype Completion Score
- What exists in papers, standards, internal designs, prototypes, partial systems, or domain-specific implementations.
- Integrated Stack Completion Score
- Whether all layers work together as a unified AI safety architecture.
The key conclusion is:
- Frontier AI safety is not at zero. But the complete AI safety stack is still early. Model-level safety and governance are the most developed. Network trust, civilizational synchronization, and long-horizon kernel continuity remain the least operationalized.
Internal Self-Audit โ GPT-5.5 Scoring
May 27, 2026 ยท Single-Agent Review Pass
- ๐งญ Instructional design: 9.2 / 10
- โ๏ธ Writing quality: 8.8 / 10
- ๐งช Technical rigor: 8.6 / 10
- ๐งฎ Benchmark usefulness: 9.2 / 10
- ๐ Narrative strength: 9.1 / 10
- ๐งฑ Visual clarity: 9 / 10
- ๐ Citation readiness: 7.5 / 10
- ๐ฅ Landmark potential: 9.3 / 10