SOTA6: 80% win rate against Claude Opus on defense & national security intelligence
Our sixth evaluation round introduces a new domain: defense, national security, and geopolitical intelligence. 20 questions spanning military procurement, strategic deterrence, critical infrastructure, and adversarial state behavior — questions where getting it wrong has consequences.
o-machine 16 – 4 Claude Opus.
o-machine 14 – 6 Gemini Pro.
This is our strongest result to date against Opus (80% win rate), extending from the autonomous driving and computational economics domains into territory where behavioral inference — reading intent from actions rather than statements — matters most.
The evaluation uses a blind 3-judge panel scoring on analytical depth, evidential grounding, and causal coherence. Panel agreement rate: 85-88%. The system processes 180,000+ behavioral signals across multiple verticals, with sub-millisecond inference entirely on CPUs.
Aggregate performance across all evaluation rounds: 71% vs Claude Opus, 62% vs Gemini Pro (1,350+ total pairwise votes).