Compare / OpenAI dots vs Instinct

OpenAI dots vs Instinct

OpenAI dots is a persistent personal agent from OpenAI (United States); Instinct is an Universal Personal Execution Agent from Instinct (Spear Street Technology) (United States). On Assistant Benchmark (1–10, updated 8 Oct 2026), OpenAI dots scores 8.4 over 7 dimensions and Instinct scores 8.5 over 12; both were scored on 7 of the same dimensions, shown below. PAEval holds 31 capability records for OpenAI dots and 15 for Instinct (as of 9 Oct 2026).

Scores stay on each evaluator's own scale and are never combined. PAEval does not pick a winner.

OpenAI dotsInstinct
VendorOpenAIInstinct (Spear Street Technology)
HQ regionUnited StatesUnited States
TypePersistent personal agentUniversal Personal Execution Agent
Pricing (as documented)The first dot is included in Pro, with no independent surcharge; Pro is currently 100/200/500 US dollars per month, and 500 includes…Not compiled
Assistant Benchmark1–10 mean of scored dimensions8.4coverage 7/158.5coverage 12/15
micro1 PersonalAgentBench% of 10 workflows with trusted completionNot scored40%95% CI 17–69%
SuperCLUE-XClaw0–100 weighted total of 5 dimensionsNot scoredNot scored
Autonomy & background workDocumentedIndependently observed
Execution environmentsDocumentedDocumented
Calls, messages & identityLimited or conditionalDocumented
Connectors & toolsLimited or conditionalDocumented
Memory & personalizationDocumentedDocumented
Channels & interfacesDocumentedDocumented
Research & deliverablesLimited or conditionalNot disclosed
Transactions & bookingsLimited or conditionalNo record
Collaboration & parallelismLimited or conditionalPartly disclosed / partly tested
Reliability & recoveryDocumentedPartly disclosed / partly tested
Permissions, security & dataDocumentedDocumented
Capability records31 (0 independent)15 (3 independent)
Capability cells show the strongest documented or observed status in each domain; availability is not task success.Add more products →

Assistant Benchmark, dimension by dimension

OpenAI dotsInstinct
Carrying out an online task9test8observed in use
Travel bookingNot tested10observed in use
Recommendation quality9test9observed in use
Purchasing a product6observed in use9observed in use
Responding to emails10test9observed in use
Proactive behavior8observed in use8observed in use
Running a routineNot tested10test
Third-party integrationsNot tested8observed in use
Permissions & privacy8observed in use5observed in use
Memory9observed in use9observed in use
Phone callsn/aNot tested
Multiplayer / groupsNot tested9unlabelled
Chained tasksNot testedNot tested
Proactive restraintNot tested8observed in use
Content creation / gamesNot testedNot tested
Source scores on the 1–10 anchor scale. "Not tested" is missing evidence, not zero.Source ↗