Grok Bot vs Instinct
Grok Bot is an Individual/Team Agent from xAI (via Cursor) (United States); Instinct is an Universal Personal Execution Agent from Instinct (Spear Street Technology) (United States). On Assistant Benchmark (1–10, updated 8 Oct 2026), Grok Bot scores 7.3 over 7 dimensions and Instinct scores 8.5 over 12; both were scored on 5 of the same dimensions, shown below. In micro1 PersonalAgentBench, Grok Bot completed 3 of 10 workflows with trusted completion and Instinct 4 of 10; the 95% intervals (11–60% and 17–69%) overlap, so the difference is not statistically separated. PAEval holds 33 capability records for Grok Bot and 15 for Instinct (as of 9 Oct 2026).
Scores stay on each evaluator's own scale and are never combined. PAEval does not pick a winner.
| Grok Bot | Instinct | |
|---|---|---|
| Vendor | xAI (via Cursor) | Instinct (Spear Street Technology) |
| HQ region | United States | United States |
| Type | Individual/Team Agent | Universal Personal Execution Agent |
| Pricing (as documented) | Cursor Pro/Pro+/Ultra are priced at US$20/60/200 per month respectively; Teams Standard/Premium are US$40/120 per user/month respectively,… | Not compiled |
| Assistant Benchmark1–10 mean of scored dimensions | 7.3coverage 7/15 | 8.5coverage 12/15 |
| micro1 PersonalAgentBench% of 10 workflows with trusted completion | 30%95% CI 11–60% | 40%95% CI 17–69% |
| SuperCLUE-XClaw0–100 weighted total of 5 dimensions | Not scored | Not scored |
| Autonomy & background work | Limited or conditional | Independently observed |
| Execution environments | Documented | Documented |
| Calls, messages & identity | Limited or conditional | Documented |
| Connectors & tools | Limited or conditional | Documented |
| Memory & personalization | Documented | Documented |
| Channels & interfaces | Documented | Documented |
| Research & deliverables | Documented | Not disclosed |
| Transactions & bookings | Limited or conditional | No record |
| Collaboration & parallelism | Limited or conditional | Partly disclosed / partly tested |
| Reliability & recovery | Limited or conditional | Partly disclosed / partly tested |
| Permissions, security & data | Limited or conditional | Documented |
| Capability records | 33 (0 independent) | 15 (3 independent) |
Assistant Benchmark, dimension by dimension
| Grok Bot | Instinct | |
|---|---|---|
| Carrying out an online task | 8observed in use | 8observed in use |
| Travel booking | Not tested | 10observed in use |
| Recommendation quality | 7observed in use | 9observed in use |
| Purchasing a product | 8observed in use | 9observed in use |
| Responding to emails | 9observed in use | 9observed in use |
| Proactive behavior | Not tested | 8observed in use |
| Running a routine | Not tested | 10test |
| Third-party integrations | 3observed in use | 8observed in use |
| Permissions & privacy | Not tested | 5observed in use |
| Memory | Not tested | 9observed in use |
| Phone calls | Not tested | Not tested |
| Multiplayer / groups | Not tested | 9unlabelled |
| Chained tasks | 8observed in use | Not tested |
| Proactive restraint | Not tested | 8observed in use |
| Content creation / games | 8observed in use | Not tested |