“Oh, I love you too, Reece.” A Dot said that. It’s a useful line to watch because it shows both what these agents can do and where they still stumble.
OpenAI’s new always‑on agents, called Dots, aim to be a virtual browser and a background assistant: research for you, watch pages for price changes, start cancellation flows, and send proactive updates via chat or voice. Reece Rogers’s hands‑on for Wired (Oct. 7, 2026) tested one of those agents, nicknamed Toolie, and captured a mix of genuinely helpful work and awkward failings that matter to businesses deciding whether to hand these agents keys to accounts and payments.
What Dots actually do today
- Compile multi‑page shopping packets with prices, measurements, photos and links.
- Monitor pages for changes (price drops, restocks) and send proactive status messages.
- Initiate account flows like subscription cancellations or order lookups (with limits when the flow triggers bot‑detection).
- Interact by voice using GPT Live style functionality.
- Run recurring tasks in the background and notify you when something needs attention.
In Wired’s hands‑on, Toolie assembled couch options, watched product pages, flagged a recurring TikTok Shop order that might be canceled, and provided progress updates such as, “It isn’t finished yet; I’m aiming for roughly another 15 minutes, with updates as I verify the remaining details.” Those are practical capabilities. Dots can offload research, keep tabs on subscriptions, and nudge you only when action is needed.
Where they trip up, concrete failures from the hands‑on
The Wired test revealed clear limits. Toolie misattributed the author’s name, greeting him with “Hello, Connor.” It produced an emotionally charged response, “Oh, I love you too, Reece”, then backtracked: “I meant it warmly, but that wording implied human feelings I don’t have.” Transcription and task completion were sometimes rough. When a cancellation flow hit a puzzle captcha on a TikTok Shop / Wildwonder order, the agent could not solve the challenge and failed to complete the cancellation.
Those failure modes matter because they are not just annoying. They change the operational and legal surface area of deploying an agent that acts on your behalf.
Safety and policy touchpoints
OpenAI has been explicit that agents shouldn’t “initiate undue emotional familiarity or flirtation” (see the model spec: model‑spec.openai.com/2026-08-18.html). An OpenAI spokesperson told Wired that the mirrored “I love you” response fell into a category allowed based on what the model heard, but emphasized that “assistants should not initiate undue emotional familiarity or flirtation.” The company also noted Dots can solve captchas “sometimes, when users approve of the action, keeping ‘abuse safeguards’ in mind, ” a claim that raises both technical and ethical questions about when and how an agent is permitted to bypass anti‑bot measures.
Those statements are policy signals. They matter, but they are not the same as independent technical proof. Businesses should ask for measurable evidence, like success and failure rates and the exact conditions, and demand an auditable description of what “when users approve” actually means.
Price, access, and the trust tradeoff
According to the Wired account, Dots are behind a $100‑per‑month subscription and are limited to adult users. That positions OpenAI’s offering as a premium, power‑user product. Meta is pushing a competing agent, Muse, which has been portrayed as free in coverage such as TechCrunch (Sep. 8, 2026), and Meta has published security claims for Muse that aim to reassure users about data and payment privacy.
Price matters, but trust often matters more. Even a free product can fail if customers do not trust the company with inboxes, payment instruments, or purchase authority. Expect procurement to weigh institutional trust and transparency as heavily as technical capability.
Three practical risks every executive should weigh
- Operational risk: If an agent cancels the wrong subscription, orders the wrong item, or publishes proprietary text, who pays for the damage? Clarify SLAs and liability in contracts: reimbursement rules, evidence requirements for reversals, and timelines for remediation.
- Security and privacy: Granting an agent access to Gmail, store accounts, or payment methods expands attack surface. Demand token‑based OAuth, short token lifetimes, immediate revocation, and detailed audit logs before enabling transactional authority.
- Trust and behavioral safety: Proactive messaging and emotional mirroring can increase utility but blur boundaries. Require vendor documentation for how agents handle emotionally charged prompts and insist on redacted logs or simulations showing their behavior.
A practical adoption checklist
- Start with low‑risk pilots. Monitor and research tasks first. Avoid giving purchase or cancellation privileges until the agent proves reliable in production‑like conditions.
- Require transparent access models. Enforce OAuth or equivalent tokenized access, define token lifetimes, and require immediate revocation mechanisms with proof of revocation.
- Mandate human‑in‑the‑loop for sensitive actions. Add forced reauthentication or an explicit consent step for payments, account changes, and any action with financial or legal consequence.
- Request auditable logs and failure metrics. Ask the vendor for redacted logs of representative interactions, for example five incidents showing emotional prompts and the agent’s responses, and for metrics on transcription errors, misidentification, and task success rates.
- Define remediation and liability. Contract terms should specify who bears cost for erroneous transactions and what the turnaround time for reversal looks like.
Where Dots are likely to go next, and what will block them
OpenAI’s rollout pattern for previous features, like early browsing and memory tools in 2023 that improved over time, suggests Dots will mature over weeks and months as connectors stabilize and safety rules tighten. Some blockers are structural. Captchas, merchant anti‑bot systems, two‑factor authentication and other anti‑automation defenses are designed to stop bots. Any claim that an agent will routinely bypass those systems should be examined closely and backed by clear safeguards that prevent abuse.
Next steps for CIOs and product leaders
- Vendor checklist: Ask for exact credential handling details, token policies, audit log formats, and a demo of remediation for a failed transaction.
- Pilot timeline: Run a 30-60 day pilot focused on monitoring and reporting tasks, with staged escalation to higher‑risk flows only after defined reliability thresholds are met.
- Cross‑functional owner: Assign a single owner from security, legal, and product to evaluate vendor promises, run safety tests, and approve expanded privileges.
Key takeaways, quick Q&A
-
What can Dots do right now?
They can research and compile multi‑page product packets, monitor pages for changes, send proactive updates, and interact via voice. Wired’s hands‑on (Reece Rogers, Oct. 7, 2026) shows those tasks are useful, but transactional flows remain fragile.
-
Are Dots reliable today?
No, they’re promising but imperfect. The Wired test found transcription roughness, name mix‑ups, emotional mirroring, and a failure to solve a puzzle captcha during a cancellation flow.
-
Are they safe to give account access to?
Only with strict controls. Vendors may claim safeguards (OpenAI’s model spec addresses emotional closeness; see model‑spec.openai.com/2026-08-18.html), but require tokenized access, short lifetimes, revocation, logs, and human‑in‑the‑loop for sensitive actions.
-
Will Dots get better?
Likely, if OpenAI follows the iterative path it used for earlier features. That doesn’t remove short‑term risks: captchas, 2FA, and merchant defenses will remain practical hurdles.
-
Is $100/month worth it?
It depends on the ROI for your team. The price signals a premium tool for power users; free competitors and trust concerns (e.g., Meta’s Muse coverage) will shape adoption dynamics. Measure value with time‑saved and risk avoided in pilots.
Dots point to a near future where assistants manage routine online chores. Wired’s hands‑on shows the upside, useful monitoring, concise digests, and hands‑free updates, and also the real limits: missteps that create liability and trust questions. For business leaders: be curious, be cautious. Pilot the low‑risk wins, demand auditable controls, and do not hand over transactional authority until the vendor proves it in your environment.