AI Voice for Email Approvals: Use as Narrator, Not Sign-Off

I Let an AI Voice Approve My Emails. 3 Were Wrong.

Siraj Raval ran a tight, uncomfortable experiment: he layered a text-to-speech voice over an existing AI agent that manages his inbox, closed the visual log, and approved every action by listening only. After the session he opened the log and announced the count: “I said I’d publish the number before I knew what it was. It’s 24 correct, 3 wrong.” (Siraj Raval, video “I Let an AI Voice Approve My Emails. 3 Were Wrong.”, 0:00).

How the test was run (method)

  • The agent’s capabilities, as described by Siraj: read incoming messages, draft replies, flag items that need human attention, and update the calendar.
  • Siraj added a single TTS layer so the agent would speak its decisions instead of showing text. He used Fish Audio S2.1 Pro and discloses a paid integration. The video description also notes, “The Developer API is free through the end of November, subject to their Fair Use Policy.”
  • He closed the text log (hid the textual record) and approved actions by listening only. The approvals were performed by Siraj himself during a single session.
  • Outcome reported: 27 approvals in total, “24 correct, 3 wrong.” (Siraj Raval, same video).

What the numbers mean, and what they don’t

This is a small, single-person, single-session experiment (n = 27). In that narrow context, the voice workflow seemed faster and produced mostly accurate outcomes. It also let three incorrect actions slip through when approval relied on audio alone. That can be consequential depending on the type of action.

Siraj sums the human factor plainly:

“give your agent a voice for output, never for approval.”, Siraj Raval (video, 3:48)

“Voice is a great narrator and a terrible auditor.”, Siraj Raval (video, 3:48)

People respond to fluent, confident speech. A confident delivery lowers skepticism and can hide content errors, a bias related to illusory-truth and narrative-persuasion effects. In short: a pleasant, authoritative voice can make false or sloppy content feel right. That is why audio works for summaries but is risky as the only audit method.

Concrete guardrails for business teams

If you’re building voice-enabled agents for email, sales outreach, customer support, or calendar automation, apply these practical rules and UI patterns:

  • Voice = narrator, not sign-off. Let the agent speak summaries and suggested drafts. Require a visible text confirmation (the exact message body) before the system sends anything that creates obligations.
  • Tier by risk with examples.
    • Informational (low risk): read-only summaries, meeting reminders, OK for voice-first review.
    • Scheduling (medium risk): proposed meeting times, require a one-click visual reveal of the exact calendar invite before confirming.
    • Legal/financial (high risk): contract language, billing changes, refunds, always require full textual review and explicit human sign-off.
  • Dual-modality approval workflow (UI pattern). Play a short audio summary, then show the collapsed draft. The approver clicks to expand the exact text and clicks a labeled “Approve message” button that records the action and preserves the text in the audit log.
  • Instrument model confidence and escalate. Surface confidence scores and explain what they mean. During pilots, escalate any item below your calibrated threshold to a human reviewer. Start conservatively and tune thresholds based on measured error severity.
  • Keep an auditable text log. Audio helps speed comprehension; the text log is the forensic record you must keep for compliance and post‑mortem diagnosis.
  • Capture and analyze errors. When voice-approved actions are wrong, capture full transcripts and metadata (time, model version, confidence, downstream effects) so you can root-cause whether the problem was prompt, model hallucination, or incorrect business rules.

Specific metrics to run during a pilot

  • Approval throughput: average seconds per approval (audio-only vs. audio+visual).
  • Error rate: wrong approvals per 100 approvals, broken down by severity (harmless, recoverable, critical).
  • Time-to-detection: how long until a wrong approval is noticed and corrected.
  • False escalation rate: percent of safe items escalated unnecessarily.

Collect those KPIs for at least several hundred approvals before you generalize the results beyond your pilot cohort. Siraj’s session is a useful nudge, not a statistical proof.

What this test leaves open, checklist to investigate

  • Severity of the three wrong approvals: Were they harmless typos, scheduling mistakes, or binding commitments?
  • Failure mode: were the errors caused by prompt design, model hallucination, mistaken intent detection, or incorrect business rules?
  • Visibility and controls: exactly how was the text log hidden and what did “approve” look like in practice (explicit command, passive acceptance)?
  • Generality: would the same error rate appear for other users, languages, or higher volumes?
  • Vendor terms and privacy: does your TTS provider’s API and paid integration meet your data residency, retention, and compliance requirements? (Siraj discloses the paid integration; see the video description for his note about the Fish Audio Developer API at time of posting.)

How to run a tight pilot that scales Siraj’s insight into safe practice

  • Start on low-risk data: route only informational threads to voice-only review for a defined pilot group.
  • Log everything: audio, the text the agent would have sent, the user’s explicit approval action, model confidence, and timestamps.
  • Define severity categories and thresholds up front. For instance, mark any item that schedules meetings with external vendors or changes billing as high severity and require text sign-off.
  • Measure and iterate weekly: error rate, severity, throughput, and time-to-detection. Tune the approval thresholds and expand scope only when metrics and manual audits justify it.
  • Train people: make “Did you read it?” an explicit cultural checkpoint. Teach reviewers to distrust confident-sounding narration and to verify the exact text for non-trivial actions.

Key takeaways, quick questions you might be asking

  • Can a voice speed up approvals?

    Yes. In Siraj’s single-session test (n = 27), he approved actions by listening alone and reported “24 correct, 3 wrong, ” indicating faster throughput but not error-free behavior.

  • Is audio reliable enough to replace a text audit?

    No. A confident voice lowers scrutiny and can hide content errors. Use spoken summaries for speed and clarity, but keep text-based approval and logs for final sign-off and auditing.

  • Which TTS did he use and what should I check?

    He used Fish Audio S2.1 Pro and disclosed a paid integration; the video description noted the Fish Audio Developer API was free through the end of November (subject to their Fair Use Policy) at time of posting. Always verify current vendor terms, privacy, and compliance before integration.

  • What should I let a voice approve?

    Allow voice-only approval only for clearly low-risk informational items. Anything that binds the organization legally, financially, or reputationally should require explicit textual review and human sign-off.

“it talks like a log file, so I skim it, and skimming isn’t checking.”, Siraj Raval (video, 3:48)

Pithy and useful: voices sell trust. Use that trust to surface clarity, not to outsource judgment. Before you let an agent RSVP, bill, or sign, decide whether you want your approval to be persuasive narration, or a deliberate, documented decision.