Compare / Town vs Ollie

Town vs Ollie

Town is a work-based personal agents and team automation from Town.com (United States); Ollie is a home personal assistant from Confabulation Corporation (United States). On Assistant Benchmark (1–10, updated 8 Oct 2026), Town scores 8.3 over 7 dimensions and Ollie scores 8.3 over 13; both were scored on 6 of the same dimensions, shown below. PAEval holds 15 capability records for Town and 16 for Ollie (as of 9 Oct 2026).

Scores stay on each evaluator's own scale and are never combined. PAEval does not pick a winner.

TownOllie
VendorTown.comConfabulation Corporation
HQ regionUnited StatesUnited States
TypeWork-based personal agents and team automationHome personal assistant
Pricing (as documented)Not compiledPublic index price: Free 6 items per day + 50 items per month; Everyday $25/month, 10 per day + 150 per month; Always-On $100/month, 15…
Assistant Benchmark1–10 mean of scored dimensions8.3coverage 7/158.3coverage 13/15
micro1 PersonalAgentBench% of 10 workflows with trusted completionNot scoredNot scored
SuperCLUE-XClaw0–100 weighted total of 5 dimensionsNot scoredNot scored
Autonomy & background workDocumentedIndependently observed
Execution environmentsPartly disclosed / partly testedPartly disclosed / partly tested
Calls, messages & identityDocumentedPartly disclosed / partly tested
Connectors & toolsDocumentedDocumented
Memory & personalizationDocumentedDocumented
Channels & interfacesDocumentedDocumented
Research & deliverablesDocumentedPartly disclosed / partly tested
Transactions & bookingsNo recordNo record
Collaboration & parallelismDocumentedDocumented
Reliability & recoveryDocumentedNot disclosed
Permissions, security & dataDocumentedDocumented
Capability records15 (1 independent)16 (2 independent)
Capability cells show the strongest documented or observed status in each domain; availability is not task success.Add more products →

Assistant Benchmark, dimension by dimension

TownOllie
Carrying out an online taskNot tested9test
Travel bookingNot tested7test
Recommendation quality8test7test
Purchasing a productNot tested8test
Responding to emails8test8test
Proactive behaviorNot tested8test
Running a routine7test10test
Third-party integrations10observed in use7test
Permissions & privacy8observed in use8test
Memory8unlabelledNot tested
Phone callsNot testedn/a
Multiplayer / groupsNot tested9test
Chained tasksNot tested8test
Proactive restraintNot tested10test
Content creation / games9test9test
Source scores on the 1–10 anchor scale. "Not tested" is missing evidence, not zero.Source ↗