Compare / Instinct vs Shuffle
Instinct vs Shuffle
Instinct is an Universal Personal Execution Agent from Instinct (Spear Street Technology) (United States); Shuffle is a creation/entertainment companion assistant, including website execution from DNFT (United States). On Assistant Benchmark (1–10, updated 8 Oct 2026), Instinct scores 8.5 over 12 dimensions and Shuffle scores 7.8 over 10; both were scored on 8 of the same dimensions, shown below. PAEval holds 15 capability records for Instinct and 14 for Shuffle (as of 9 Oct 2026).
Scores stay on each evaluator's own scale and are never combined. PAEval does not pick a winner.
| Instinct | Shuffle | |
|---|---|---|
| Vendor | Instinct (Spear Street Technology) | DNFT |
| HQ region | United States | United States |
| Type | Universal Personal Execution Agent | Creation/entertainment companion assistant, including website execution |
| Pricing (as documented) | Not compiled | Pro $8/month: Unlimited chat + 880 monthly picture/video/music credits; $1 daily pass and $4.99 weekly pass are unlimited chat. |
| Assistant Benchmark1–10 mean of scored dimensions | 8.5coverage 12/15 | 7.8coverage 10/15 |
| micro1 PersonalAgentBench% of 10 workflows with trusted completion | 40%95% CI 17–69% | Not scored |
| SuperCLUE-XClaw0–100 weighted total of 5 dimensions | Not scored | Not scored |
| Autonomy & background work | Independently observed | Partly disclosed / partly tested |
| Execution environments | Documented | Documented |
| Calls, messages & identity | Documented | Partly disclosed / partly tested |
| Connectors & tools | Documented | Not disclosed |
| Memory & personalization | Documented | Documented |
| Channels & interfaces | Documented | Documented |
| Research & deliverables | Not disclosed | Documented |
| Transactions & bookings | No record | No record |
| Collaboration & parallelism | Partly disclosed / partly tested | Documented |
| Reliability & recovery | Partly disclosed / partly tested | Independently observed |
| Permissions, security & data | Documented | Documented |
| Capability records | 15 (3 independent) | 14 (1 independent) |
Capability cells show the strongest documented or observed status in each domain; availability is not task success.Add more products →
Assistant Benchmark, dimension by dimension
| Instinct | Shuffle | |
|---|---|---|
| Carrying out an online task | 8observed in use | 7.5test |
| Travel booking | 10observed in use | 9test |
| Recommendation quality | 9observed in use | 9test |
| Purchasing a product | 9observed in use | 8test |
| Responding to emails | 9observed in use | 8test |
| Proactive behavior | 8observed in use | Not tested |
| Running a routine | 10test | Not tested |
| Third-party integrations | 8observed in use | 7test |
| Permissions & privacy | 5observed in use | 7test |
| Memory | 9observed in use | 6test |
| Phone calls | Not tested | n/a |
| Multiplayer / groups | 9unlabelled | n/a |
| Chained tasks | Not tested | 7test |
| Proactive restraint | 8observed in use | Not tested |
| Content creation / games | Not tested | 9test |
Source scores on the 1–10 anchor scale. "Not tested" is missing evidence, not zero.Source ↗