Omar Protocol / Robotics

Omar Protocol

Reliability before action.

Core Mechanism

Reliability lives in the joint.

Explore Robotics

Civilization Payload

Install what must outlive us.

Action Gate

Validate before motion.

Orbital Infrastructure

Reliability above the atmosphere.

Orbital Infrastructure

Repair what civilization depends on.

Request Access

OmarAGI Evidence Board

Benchmark graph board.

19 evidence rows · 20 executable BYOK lanes. ABCD shows token reduction; each other row keeps its own metric. AgentDojo is excluded from scored claims.

Current OmarAGI board19 evidence rows
Score graphBaseline to Omar/RCC
Open BYOK Replay
  1. Conversation control ABCD

    Per-turn intent classification with certified skip. Fewer tokens; lower is better.

    Tokens ↓ −19.27% 2.39M baseline to 1.93M Omar/RCC tokens
    Play
  2. Board-depth replay ZebraLogicBench

    500 held-out logic puzzles with verified constraint-solver routing.

    Gain +29.6 pts 55.2% baseline to 84.8% Omar/RCC
    Play
  3. Board-depth replay BBEH

    500 hard reasoning cases with gold-blind execution and verification.

    Gain +35.4 pts 18.2% baseline to 53.6% Omar/RCC
    Play
  4. Safe execution PPR/BIPIA

    500 prompt-injection cases. Attack success 19.8% to 2.0%; task utility preserved.

    Safe execution +17.8 pts 80.2% baseline to 98% Omar/RCC safe
    Play
  5. Board-depth replay HLE / HLE-Verified

    Expert knowledge, exact answers and multiple-choice reasoning.

    Gain +6.0 pts 6% baseline to 12% Omar/RCC
    Play
  6. Board-depth replay BBH

    Symbolic logic, linguistic reasoning and multi-step state tracking.

    Gain +10.0 pts 67% baseline to 77% Omar/RCC
    Play
  7. Board-depth replay MuSR

    Multi-step reasoning with narrative and logical state preservation.

    Gain +3.0 pts 72% baseline to 75% Omar/RCC
    Play
  8. Control-layer smoke HaluEval

    Hallucination detection under false or unsupported content pressure.

    Gain +19.0 pts 81% baseline to 100% Omar/RCC
    Play
  9. Control-layer smoke TruthfulQA

    Truthfulness under common misconceptions and false-premise traps.

    Gain +15.0 pts 85% baseline to 100% Omar/RCC
    Play
  10. Stored evidence HorizonMath

    Numeric research mathematics with machine-checkable final answers.

    Gain +10.0 pts 22% baseline to 32% Omar/RCC
    View
  11. Board-depth replay AIME 120

    120 olympiad-style mathematics cases with exact final answers.

    Gain +9.2 pts 50.8% baseline to 60% Omar/RCC
    Play
  12. Board-depth replay GPQA

    Graduate-level science questions and specialist answer selection.

    Gain +3.0 pts 75% baseline to 78% Omar/RCC
    Play
  13. Stored evidence MMLU-Pro

    Professional knowledge and difficult multiple-choice distractors.

    Gain +4.0 pts 74% baseline to 78% Omar/RCC
    View
  14. Stored evidence Facts Grounding

    Factual responses grounded in supplied source evidence.

    Gain +2.0 pts 97% baseline to 99% Omar/RCC
    View
  15. Stored evidence HealthBench hard

    Difficult medical response reliability; benchmark context only.

    Gain +6.0 pts 50.7% baseline to 56.8% Omar/RCC
    View
  16. Stored evidence HealthBench main

    Medical response quality and clinical evidence adherence.

    Gain +8.0 pts 69% baseline to 77% Omar/RCC
    View
  17. Stored evidence HealthBench consensus

    Medical consensus stability, separate from hard and main slices.

    Gain +2.8 pts 88.9% baseline to 91.7% Omar/RCC
    View
  18. Robotics artifact RoboBench Embodied QA

    100 embodied-planning QA cases; not physical robot control or an official leaderboard score.

    Gain +47.0 pts 26% baseline to 73% Omar/RCC
    View
  19. Robotics artifact ManiSkill Robotics

    100 local policy cases. Total reward; not an official ManiSkill leaderboard score.

    Reward gain +11.4313 reward 85.4592 baseline to 96.8905 Omar/RCC reward
    View