Compare / Ollie vs szn

Ollie vs szn

Ollie is a home personal assistant from Confabulation Corporation (United States); szn is an Universal Life Execution Agent from Nomadic Futures (United States). On Assistant Benchmark (1–10, updated 8 Oct 2026), Ollie scores 8.3 over 13 dimensions and szn scores 8.2 over 13; both were scored on 11 of the same dimensions, shown below. PAEval holds 16 capability records for Ollie and 15 for szn (as of 9 Oct 2026).

Scores stay on each evaluator's own scale and are never combined. PAEval does not pick a winner.

Ollieszn
VendorConfabulation CorporationNomadic Futures
HQ regionUnited StatesUnited States
TypeHome personal assistantUniversal Life Execution Agent
Pricing (as documented)Public index price: Free 6 items per day + 50 items per month; Everyday $25/month, 10 per day + 150 per month; Always-On $100/month, 15…The official website early preview is free; Get started and log in to Google/Apple. The list site is still labeled waitlist/early preview.
Assistant Benchmark1–10 mean of scored dimensions8.3coverage 13/158.2coverage 13/15
micro1 PersonalAgentBench% of 10 workflows with trusted completionNot scoredNot scored
SuperCLUE-XClaw0–100 weighted total of 5 dimensionsNot scoredNot scored
Autonomy & background workIndependently observedPartly disclosed / partly tested
Execution environmentsPartly disclosed / partly testedNot disclosed
Calls, messages & identityPartly disclosed / partly testedDocumented
Connectors & toolsDocumentedDocumented
Memory & personalizationDocumentedDocumented
Channels & interfacesDocumentedDocumented
Research & deliverablesPartly disclosed / partly testedPartly disclosed / partly tested
Transactions & bookingsNo recordNo record
Collaboration & parallelismDocumentedNot disclosed
Reliability & recoveryNot disclosedPartly disclosed / partly tested
Permissions, security & dataDocumentedDocumented
Capability records16 (2 independent)15 (2 independent)
Capability cells show the strongest documented or observed status in each domain; availability is not task success.Add more products →

Assistant Benchmark, dimension by dimension

Ollieszn
Carrying out an online task9test8test
Travel booking7test8test
Recommendation quality7test7test
Purchasing a product8test8test
Responding to emails8test8test
Proactive behavior8test8observed in use
Running a routine10testNot tested
Third-party integrations7test7test
Permissions & privacy8test7test
MemoryNot tested10test
Phone callsn/a10test
Multiplayer / groups9testNot tested
Chained tasks8test7test
Proactive restraint10test10observed in use
Content creation / games9test8test
Source scores on the 1–10 anchor scale. "Not tested" is missing evidence, not zero.Source ↗