Compare / Tomo vs Caddy
Tomo vs Caddy
Tomo is a personal goals/life companionship and execution assistant from Mapo Labs (United States); Caddy is a Message-Based Universal Personal/Home Execution Agent from Caddy Technologies Inc. (United States). On Assistant Benchmark (1–10, updated 8 Oct 2026), Tomo scores 7.7 over 11 dimensions and Caddy scores 7.3 over 9; both were scored on 9 of the same dimensions, shown below. PAEval holds 14 capability records for Tomo and 15 for Caddy (as of 9 Oct 2026).
Scores stay on each evaluator's own scale and are never combined. PAEval does not pick a winner.
| Tomo | Caddy | |
|---|---|---|
| Vendor | Mapo Labs | Caddy Technologies Inc. |
| HQ region | United States | United States |
| Type | Personal goals/life companionship and execution assistant | Message-Based Universal Personal/Home Execution Agent |
| Pricing (as documented) | The official website base starts at $19.99/month, other plans increase the message limit; the price can be different due to… | Public beta, free to start; list site records Free public beta, no standard payment price list. |
| Assistant Benchmark1–10 mean of scored dimensions | 7.7coverage 11/15 | 7.3coverage 9/15 |
| micro1 PersonalAgentBench% of 10 workflows with trusted completion | Not scored | Not scored |
| SuperCLUE-XClaw0–100 weighted total of 5 dimensions | Not scored | Not scored |
| Autonomy & background work | Partly disclosed / partly tested | Documented |
| Execution environments | Not disclosed | Partly disclosed / partly tested |
| Calls, messages & identity | Documented | Documented |
| Connectors & tools | Documented | Documented |
| Memory & personalization | Documented | Partly disclosed / partly tested |
| Channels & interfaces | Documented | Documented |
| Research & deliverables | Partly disclosed / partly tested | Partly disclosed / partly tested |
| Transactions & bookings | No record | No record |
| Collaboration & parallelism | Documented | Documented |
| Reliability & recovery | Partly disclosed / partly tested | Partly disclosed / partly tested |
| Permissions, security & data | Documented | Documented |
| Capability records | 14 (1 independent) | 15 (1 independent) |
Capability cells show the strongest documented or observed status in each domain; availability is not task success.Add more products →
Assistant Benchmark, dimension by dimension
| Tomo | Caddy | |
|---|---|---|
| Carrying out an online task | 7test | 8test |
| Travel booking | 9test | 8test |
| Recommendation quality | 8test | 7test |
| Purchasing a product | 8test | 7test |
| Responding to emails | 8test | 7test |
| Proactive behavior | Not tested | Not tested |
| Running a routine | Not tested | Not tested |
| Third-party integrations | 7test | 7test |
| Permissions & privacy | 7test | Not tested |
| Memory | 6test | 7test |
| Phone calls | Not tested | n/a |
| Multiplayer / groups | 9observed in use | n/a |
| Chained tasks | 7test | 7test |
| Proactive restraint | Not tested | Not tested |
| Content creation / games | 9test | 8test |
Source scores on the 1–10 anchor scale. "Not tested" is missing evidence, not zero.Source ↗