Compare / Grok Bot vs szn

Grok Bot vs szn

Grok Bot is an Individual/Team Agent from xAI (via Cursor) (United States); szn is an Universal Life Execution Agent from Nomadic Futures (United States). On Assistant Benchmark (1–10, updated 8 Oct 2026), Grok Bot scores 7.3 over 7 dimensions and szn scores 8.2 over 13; both were scored on 7 of the same dimensions, shown below. PAEval holds 33 capability records for Grok Bot and 15 for szn (as of 9 Oct 2026).

Scores stay on each evaluator's own scale and are never combined. PAEval does not pick a winner.

Grok Botszn
VendorxAI (via Cursor)Nomadic Futures
HQ regionUnited StatesUnited States
TypeIndividual/Team AgentUniversal Life Execution Agent
Pricing (as documented)Cursor Pro/Pro+/Ultra are priced at US$20/60/200 per month respectively; Teams Standard/Premium are US$40/120 per user/month respectively,…The official website early preview is free; Get started and log in to Google/Apple. The list site is still labeled waitlist/early preview.
Assistant Benchmark1–10 mean of scored dimensions7.3coverage 7/158.2coverage 13/15
micro1 PersonalAgentBench% of 10 workflows with trusted completion30%95% CI 11–60%Not scored
SuperCLUE-XClaw0–100 weighted total of 5 dimensionsNot scoredNot scored
Autonomy & background workLimited or conditionalPartly disclosed / partly tested
Execution environmentsDocumentedNot disclosed
Calls, messages & identityLimited or conditionalDocumented
Connectors & toolsLimited or conditionalDocumented
Memory & personalizationDocumentedDocumented
Channels & interfacesDocumentedDocumented
Research & deliverablesDocumentedPartly disclosed / partly tested
Transactions & bookingsLimited or conditionalNo record
Collaboration & parallelismLimited or conditionalNot disclosed
Reliability & recoveryLimited or conditionalPartly disclosed / partly tested
Permissions, security & dataLimited or conditionalDocumented
Capability records33 (0 independent)15 (2 independent)
Capability cells show the strongest documented or observed status in each domain; availability is not task success.Add more products →

Assistant Benchmark, dimension by dimension

Grok Botszn
Carrying out an online task8observed in use8test
Travel bookingNot tested8test
Recommendation quality7observed in use7test
Purchasing a product8observed in use8test
Responding to emails9observed in use8test
Proactive behaviorNot tested8observed in use
Running a routineNot testedNot tested
Third-party integrations3observed in use7test
Permissions & privacyNot tested7test
MemoryNot tested10test
Phone callsNot tested10test
Multiplayer / groupsNot testedNot tested
Chained tasks8observed in use7test
Proactive restraintNot tested10observed in use
Content creation / games8observed in use8test
Source scores on the 1–10 anchor scale. "Not tested" is missing evidence, not zero.Source ↗