Grok Bot
SpaceXAI/xAI, Cursor hosting
As of 9 Oct 2026, PAEval has compiled 33 capability records for Grok Bot from 17 source URLs; 28 come from the vendor's own documentation and none has yet been independently observed. Assistant Benchmark (updated 8 Oct 2026) gives it 7.3 on its 1–10 scale, the mean of the 7 of 15 dimensions it scored, ranking #11 of 12 scored products. In micro1 PersonalAgentBench it completed 3 of 10 workflows without a critical violation (30% trusted completion; 95% interval 11–60%), and 3 of 10 counting any completion. Scores from different evaluators use different scales and are not combined. PAEval does not publish its own score or rank for it.
Current desktop/mobile Grok Bot; visible changelog v0.68.0 (10-05), unfixed background model; personal Bot, Team Bot and template classification, recording Cursor/SuperGrok access and Auto Review. Not Grok Chat, Grok Build, or the Model API.
- Independent0
- Vendor28
- Third-party report0
- Not established5
33 records · 17 source URLs · 2 external result(s)
Capability summary
Strongest status per domain; every underlying record is listed below. Availability does not establish task success.
Positioning & scope
Named Bots for repetitive work and continuous responsibilities; individual and Team Bot classification to support inter-Bot collaboration.
Regions, eligibility & access
All paid Cursor individuals and Teams include usage rights; can be associated with SuperGrok/Plus/Heavy/X Premium+ or corresponding Business seats; Lite is not included; businesses need to contact to activate.
macOS Apple/Intel, Windows x64/Arm, Linux x64/Arm, iOS/iPadOS18+, Android9+; desktop 20+ interface languages and continues to increase.
The complete list of registration/payment countries has not been verified; the cloud computer is located in the United States, which does not mean that it is limited to US users.
Pricing & quotas
Cursor Pro/Pro+/Ultra are priced at US$20/60/200 per month respectively; Teams Standard/Premium are US$40/120 per user/month respectively, tax excluded; Bot does not require independent subscription.
The weekly usage quota is measured in the Cursor account; SuperGrok and Cursor quotas do not overlap; when used up, you can enable on-demand payment and set a monthly limit. Teams enables on-demand by default.
Limitations: The number of stable tasks/exact weekly tokens for each level has not been announced; using up the weekly quota when on-demand does not mean stopping deductions.
The trial period is based on the usage credit plus a 7-day window, and there is no guarantee that the seven days will be sufficient; the account cannot be removed/transferred to other Cursor accounts after it is linked.
Autonomy & background work
Supports background timing and supported events/independent webhooks; event connections are different from ordinary plug-ins; a single Bot can have up to 50 routines, a minimum of 5 minutes, and retains the latest 20 records.
Shared skills can be saved from successful work; Teach a task generates skills to be verified with a screen demonstration of up to 10 minutes, which will be gradually opened.
Background approval initiated by non-current user chats expires in about 10 minutes and will not be executed; long-term inactivity may ask and suspend routines.
Execution environments
Each user has a Cursor cloud computer, and all Bots share files, cookies/login/CLI credentials; each Bot's independent screen can be run in parallel, and one computer operation task is performed on one screen at a time.
Mac/Windows native commands can be authorized; by default, each time you ask, you can always allow/disable, and the management policy can be stricter.
Limitations: Native execution is not equivalent to self-hosted Grok Bot; on-prem/BYO images are not officially supported.
The mobile phone is connected to the same cloud computer, and the native cloud phone/mobile GUI execution is not verified.
Calls, messages & identity
Can prepare and send emails/Slack messages; individual Bot use is granted account permissions; Team Bot can have an independent Slack app to present its identity.
Limitations: Slack's presentation identity is not equal to independent IAM permissions; OAuth uses the questioner's account, and key-type integration can use team Bot shared credentials.
The official team Marketplace lists dial bot to initiate outbound calls through Bland AI, which requires Bland API key/voice for the first time.
Conditions: External service integration/templates instead of built-in telephony network.
Limitations: Cost, number/language/region and service results are not verified.
Unverified native phone call answering, user callback, or automatic meeting participation; real-time Bot voice is not equal to PSTN outbound calls.
Connectors & tools
The Marketplace plug-in can operate and support services, and all Bots in the account are available; the Team Bot document clearly identifies Remote HTTPS and Command-type custom MCP, and distinguishes between OAuth and shared key identities.
There is no public product API for direct programmatic control of persistent Bots in this round; the model API and Grok Build MCP cannot be replaced.
Memory & personalization
Name the Bot to save your own work context/preferences, and share at the cloud file/session account level; copying a template does not copy history, learning memories or attachments.
Team Bot is divided into team memory and each person's private context; shared information should be explicitly specified; team memory can be requested to be viewed/corrected/deleted.
Personal Bot's item-by-item memory editing/deletion UI and export integrity have not been verified this round; deleting a Bot does not delete shared files/sessions.
Channels & interfaces
Synchronize with Bot across desktops and mobile devices; text, pictures/documents, voice input, real-time voice and agent voice messages.
X can be searched publicly; the @bot entrance needs to be associated with X and not all areas are open.
Research & deliverables
Input images/audio/video, Office/PDF/CSV/code/notebook; generate reviewable documents, tables, slides, etc.; the desktop can have up to 6 attachments at a time, normal 25MB, video 200MB.
The cloud terminal code can be executed and the encoding handed over to the Cloud Agent, and the final file/image/video can be sent back to the session.
Transactions & bookings
Changes, purchases, etc. are subject to approval/takeover; in August, the changelog already had a Link purchase approval card.
Limitations: Merchant coverage and order recovery/refund capabilities have not been verified.
Collaboration & parallelism
2 to 6 Bot group chats can be created and handed over; sensitive approval/login/secret requests and email Slack drafts are not displayed in the group, and handoff between Bot groups only involves text.
Team Bot is published to the team, and individual chats are private to each other; the shared configuration can be maintained by the owner/manager, and members use OAuth to use their own permissions.
Reliability & recovery
Observe the computer, retry/Recover/Update; restore to retain persistent files/login, Reset to the latest snapshot may lose recent changes, and temporary software packages are not guaranteed to be durable.
Consumption session display tool activities; enterprise Audit logs, Action Recording, and OpenTelemetry are set separately and depend on the package, and Action Recording is turned off by default.
Permissions, security & data
Auto Review is a model review; Ask first takes priority; microVM is isolated between users, but Bots are not security boundaries. Banning plug-ins does not block browser access on the same website.
According to Cursor data settings, it is not compatible with Legacy Privacy Mode; US hosting does not default to Cursor US-only residency contract; there is no self-hosting/BYO image.
Deleting the Bot does not clear the shared computer/login; the model is managed by Cursor without user picker, and the enterprise model allowlist is not guaranteed to be executed.
External measurements & hands-on results
Different methods, samples and dates. Not combined into one success rate.
One user Grok Bot/Grok4.6, Cursor Ultra, 8 role Bot shared cloud computer, 5 days of real work ↗
- Evidence type
- independent_observational_task_log
- Published
- 2026-09-02
- Sample
- tasks: 67; days: 5; testers: 1
- Reported
- completed without help: 40; completed after help: 14; partial: 5; failed or cancelled: 8; interventions: 17
- Findings
- Report an accidental X post, loss of self-installed software after cloud update, and cloud IP login obstruction.
- Limitations
- The 17 interventions and each result category do not have the same denominator, so do not add them together.; Sample self-selection, mixed task difficulty, old version; not the current success rate of all users, nor can it be directly ranked with other task sets.; There was no retest in this study.
User Windows computer usage tasks ↗
- Evidence type
- incident_report
- Published
- 2026-09-26
- Findings
- Forum replies indicate that interrupting new messages and restarting the background computer assistant may result in no progress.
- Limitations
- Single event, status/repair coverage is not verified and does not translate into failure rate.
Analyst notes
Key differences
- Multiple persistent role Bots and Bot groups are the core, but sharing computers/logins with the same user does not constitute mutual isolation.
- Weekly quota, on-demand payment and Cursor/SuperGrok identity linkage are complicated; there is an official on-demand monthly cap, and the old "no cap" comment is outdated.
- It has observable routine history, 5-minute frequency minimum and approval expiration rules, which is suitable for establishing a separate evaluation dimension for unattended failures.
- Team Bot can share configurations while retaining members’ private chats; API key and OAuth permission sources must be separated.
Known unknowns
- Full region list
- task language performance
- Native mobile execution
- Native PSTN/conference participation ability
- Persistent Bot public call API
- Personal item-by-item memory control
- Accurate total concurrency upper limit/weekly quota/long-range SLA
Source conflicts and how they were resolved
- topic: computer_allocation; older: Overview MarketingEvery Bot has a computer and is easy to misread; newer: computer/security specifies that each user shares a computer and each Bot has an independent screen.; resolution: The isolation unit is recorded as user, and the screen concurrency unit is recorded as Bot.; source ids: grok_overview; grok_computer; grok_security
- topic: on_demand_spend_cap; resolution: The old third-party no-cap statement is no longer used; the current official billing page has a monthly limit.; source ids: grok_plans
- topic: sessions_after_update; resolution: The user recovery document states that the login is retained, but the security FAQ still prompts that the session will be lost during reconstruction; the persistence design cannot be written to ensure that all website sessions will last forever.; source ids: grok_computer; grok_security
Detailed capability observations
34 capabilities in 8 groups, reviewed cell by cell. Unknown cells were not researched or lacked sufficient public evidence.
| External contact and communication | ||
|---|---|---|
| User initiates real-time voiceTwo-way real-time voice between the user and his own agent; it does not mean that the agent calls out to a third party. | Documented | Officially supports users to initiate real-time voice chat. |
| The agent proactively calls the userAgent initiates a call to the user; records in-app calls or phone network. | Unknown | Not researched yet |
| Outbound calls to third partiesDial business or personal numbers as authorized and have two-way communication; voice chat and generated audio are not… | Unknown | No official information has been found that is sufficient to confirm that the product can natively make third-party calls; the availability/unavailability of all products cannot be inferred from the in-app voice or this account tool. |
| Answer external callsA third party can dial into a number or call portal controlled by the agent, and the agent will actually answer the… | Unknown | Not researched yet |
| Send and receive messages across channelsActual sending and receiving via email, SMS/RCS, WhatsApp, Slack, Teams, etc.; only drafting, not sending. | Unknown | Not researched yet |
| Participate in live meetingsActually join the meeting and listen, transcribe or speak; recording each operation and reading the meeting minutes is… | Unknown | Not researched yet |
| Agent independent external identityAgents have independent email addresses, numbers or platform identities; they are strictly distinguished from borrowed… | Unknown | Not researched yet |
| Execution environments and devices | ||
| Cloud browser operationThe agent opens web pages, operation forms and cross-site processes in the cloud; ordinary search interfaces are not… | Documented | Officially supports cloud browser and web page operations. |
| Cloud computers and terminalsAgents have operational desktops, file systems, and/or terminals; record persistence and isolation units. | Documented | The official provides persistent cloud computers, file systems and terminals; multiple Bots with the same account can share computers. |
| Cloud Phone and Mobile AppThe agent can operate apps in cloud Android/iOS or equivalent mobile systems; mobile clients and mobile remote control… | Unknown | No evidence was found that the product natively provides a cloud phone that can operate the mobile system/App; the mobile client and cloud computer cannot be used as alternative evidence. |
| User computer operationAccess files or operate applications on authorized user computers; clarify system, online conditions and operating… | Documented | Officially supports executing commands on approved Mac or Windows machines; the GUI operation scope has not been verified this time. |
| User mobile phone operationReading and writing the system/App or operating UI on the authorized user's mobile phone; only notification and… | Unknown | Not researched yet |
| Application connections and tools | ||
| Connector readRead accounts, mail, calendars, files, etc. through supported connectors; log services and read objects. | Documented | Installing connectors from the Marketplace to access services is officially supported. |
| Connector writeCreate, modify, or send via a supported connector; having a read connector does not imply write access. | Unknown | Not researched yet |
| Custom connectors and toolsSupport users/developers to access custom APIs, MCPs or plug-ins; record protocols, hosting and permissions. | Unknown | Not researched yet |
| Login and session persistenceYou can securely authorize services and reuse the login for subsequent tasks; it is not equivalent to obtaining the… | Documented | Official Description Cloud Browser login sessions persist and are shared by Bots with the same account. |
| Continuous execution and initiative | ||
| Offline background promotionAfter closing the client, authorized tasks can still continue; the duration and amount are separately noted. | Documented | Official instructions indicate that closing the app or computer does not stop cloud work. |
| Timing and recurring tasksRun tasks by date, time zone, or recurring schedule; only calendar reminders need to be noted. | Documented | Officially supports routine scheduled running. |
| event triggered executionIt is started by events such as emails, messages, Webhooks, and file changes; scheduled polling is a separate… | Limited or conditional | Officially supports event triggering for some account integrations, such as Slack messages or GitHub notifications. |
| Proactive discovery and follow-upOpportunities, reminders or suggestions are found when no new instructions are received; independent writing requires… | Unknown | Not researched yet |
| Memory and personalization | ||
| Long-term memory and preferencesPreserve and use preferences, context, or project state across sessions; don't treat context windows as long-term… | Documented | Official description: Preserve preferences and role context across sessions. |
| Cross-channel continuationDifferent entries can continue the same task and related context; message mirroring and context can be separated. | Documented | Mobile and desktop access the same set of Bots, conversations, and cloud computers. |
| Memory review and correctionUsers can view, correct or selectively forget; distinguish between full account reset and selective deletion. | Unknown | Not researched yet |
| Achievements and realistic tasks | ||
| Research and traceabilitySearch, synthesize and present verifiable sources; source accuracy documented by external review. | Unknown | Not researched yet |
| Documentation and multimodal resultsCreate or modify available products such as documents, tables, presentations, images/audio and videos; format as facet. | Unknown | Not researched yet |
| Coding and runnable deliveryImplement, run, test code, and deliver the application or change; only generating snippets is distinguished from… | Unknown | Not researched yet |
| Shopping reservation and service processingActual execution of shopping, reservation, unsubscription or service process; stopping at draft, shopping cart,… | Unknown | Not researched yet |
| Collaboration and task closed loop | ||
| Cross-application multi-step processThe same task transfers information and executes across multiple applications; multiple independent plug-ins are not… | Documented | The official description handles multi-step work across web pages and applications. |
| Parallel tasks and agent collaborationMultiple tasks/agents advance and hand off simultaneously; logging whether files, sessions, and identities are shared. | Documented | Multiple Bots can work in parallel and hand over. |
| Failure recovery and waiting for follow-upIt can be recovered after failure, waiting for reply or external changes; the completed effect must be supported by… | Unknown | Not researched yet |
| Authorization, security and auditability | ||
| Fine-grained authorizationAllow, deny, or request confirmation by action, object, and scope; presence of controls does not equal security score. | Documented | Officially provides sequential/continuous approval and rejection rules. |
| Manual takeover and continuationIt can be handed over to the user to complete the sensitive/blocked steps, and then returned to the agent to continue. | Documented | Sensitive steps can be taken over by the user to continue. |
| Activity and result trackingDisplay tool actions, execution records, receipts, task status or sources; the display range is noted separately. | Documented | The official provides computer operation previews and routine recent operation records; it does not claim to be complete and non-tamperable audits. |
| Export, Undo and DeleteSupports data export, connection cancellation or task/data deletion; the three types of operations retain evidence… | Unknown | Not researched yet |