Compare / Instinct vs szn

Instinct vs szn

Instinct is an Universal Personal Execution Agent from Instinct (Spear Street Technology) (United States); szn is an Universal Life Execution Agent from Nomadic Futures (United States). On Assistant Benchmark (1–10, updated 8 Oct 2026), Instinct scores 8.5 over 12 dimensions and szn scores 8.2 over 13; both were scored on 10 of the same dimensions, shown below. PAEval holds 15 capability records for Instinct and 15 for szn (as of 9 Oct 2026).

Scores stay on each evaluator's own scale and are never combined. PAEval does not pick a winner.

Instinctszn
VendorInstinct (Spear Street Technology)Nomadic Futures
HQ regionUnited StatesUnited States
TypeUniversal Personal Execution AgentUniversal Life Execution Agent
Pricing (as documented)Not compiledThe official website early preview is free; Get started and log in to Google/Apple. The list site is still labeled waitlist/early preview.
Assistant Benchmark1–10 mean of scored dimensions8.5coverage 12/158.2coverage 13/15
micro1 PersonalAgentBench% of 10 workflows with trusted completion40%95% CI 17–69%Not scored
SuperCLUE-XClaw0–100 weighted total of 5 dimensionsNot scoredNot scored
Autonomy & background workIndependently observedPartly disclosed / partly tested
Execution environmentsDocumentedNot disclosed
Calls, messages & identityDocumentedDocumented
Connectors & toolsDocumentedDocumented
Memory & personalizationDocumentedDocumented
Channels & interfacesDocumentedDocumented
Research & deliverablesNot disclosedPartly disclosed / partly tested
Transactions & bookingsNo recordNo record
Collaboration & parallelismPartly disclosed / partly testedNot disclosed
Reliability & recoveryPartly disclosed / partly testedPartly disclosed / partly tested
Permissions, security & dataDocumentedDocumented
Capability records15 (3 independent)15 (2 independent)
Capability cells show the strongest documented or observed status in each domain; availability is not task success.Add more products →

Assistant Benchmark, dimension by dimension

Instinctszn
Carrying out an online task8observed in use8test
Travel booking10observed in use8test
Recommendation quality9observed in use7test
Purchasing a product9observed in use8test
Responding to emails9observed in use8test
Proactive behavior8observed in use8observed in use
Running a routine10testNot tested
Third-party integrations8observed in use7test
Permissions & privacy5observed in use7test
Memory9observed in use10test
Phone callsNot tested10test
Multiplayer / groups9unlabelledNot tested
Chained tasksNot tested7test
Proactive restraint8observed in use10observed in use
Content creation / gamesNot tested8test
Source scores on the 1–10 anchor scale. "Not tested" is missing evidence, not zero.Source ↗