AI monitoring of Multiverse instructors fuels stress and surveillance concerns

“Since all of this has come in, the stress level has been absolutely horrendous.”

The line, from an anonymous Multiverse instructor, sums up a human consequence of a technical rollout. An AI that records, transcribes and rates online apprenticeship sessions has left some teachers sleepless, constantly self-monitoring and feeling under unrelenting surveillance.

Multiverse, the UK training company co-founded by Euan Blair, has deployed the system. The company is reported to be valued at about £1.6bn, with turnover of roughly £118m and losses of £28m last year. It employs more than 800 people and trains learners for clients including the NHS, local councils, universities and private employers, according to reporting. Multiverse says the system uses a mix of models including products from Anthropic and OpenAI.

What the AI does, in plain terms

Where instructors were previously observed by a manager roughly once a month, the new agent transcribes live online sessions, scores the transcript against a rubric and sends outputs to managers: a “risk status” label, a confidence percentage and written comments. The change moves observation from occasional human review to continuous automated scoring.

Reportedly, the AI is configured to flag behaviours such as:

  • slow handling of online connection problems (the internal benchmark given was roughly a minute);
  • use of “filler” phrases and hedges (examples reported: “sort of”, “kind of”);
  • vague or uncertain answers, or prolonged one-to-one exchanges that sideline the group;
  • learners going quiet, interpreted as possible disengagement.

The internal rubric, as reported, instructs: “1-2 hedges or filler words in an otherwise direct, well-internalised explanation do NOT fail this standard … What fails the standard is a PERVASIVE REPEATED pattern of filler, hedging or re-explanation.”

The guidance reportedly tells the model to “judge the PATTERN across the whole transcript, not isolated instances”.

The rubric, again as reported, tells the system to apply a negative mark for visible self‑correction “even if it was eventually straightened out”.

How teachers describe the effect

Multiple instructors who spoke anonymously say the system has changed how they teach and how they feel at work. Common themes are higher stress, fewer spontaneous moments in class, and a persistent worry that an algorithm will mark them down instead of helping them improve.

“It has affected sleep, and not just for me. I’ve gone through a lot of things in my life, and I haven’t ever been in a position where I’m waking up in the middle of the night.”

“This is inducing stress because you consistently have to be cautious. Your attention is no longer delivering [the lesson] and listening to the learner [but] am I saying the right thing so the AI doesn’t mark me down for something stupid that I’ve mistakenly said.”

“It feels intrusive. Multiverse calls it quality assurance, but it feels like surveillance.”

Teachers also say they have little visibility into the system now: only managers can see the dashboard and the AI’s annotations, and staff say they will be given access later. That asymmetry, where machine scoring is visible to managers but not to the people being scored, fuels mistrust.

What Multiverse says

A Multiverse spokesperson argued the AI “does not replace human judgment but ‘directs human time to where it’s needed most’ by highlighting problem[s], ” and that it is “using AI to improve the experience of learners and customers.” The company said only human managers write teachers’ performance reviews and described the agent as one that “provides instructors with improved advice on development and feedback, exactly how any learning organisation should be deploying AI.” Multiverse added that “the most consequential performance management action it can take is to recommend that a human personally reviews a session.”

Ofsted’s February inspection is also part of the context: the regulator found Multiverse “failed to meet expected standards in several areas, but praised the instructors as ‘skilled, effective teachers’.”

This fits a broader pattern

Multiverse is not unique. Across sectors, employers have been trialling AI to monitor and score frontline work, from analysing customer interactions in retail and fast food to experiments inside tech firms. For example, Burger King has used automated systems to analyse customer-staff interactions, and Meta paused tracking employees’ computer keystrokes amid privacy concerns (reported June 24, 2026).

Labour bodies are pushing back. The United Tech and Allied Workers branch of the Communication Workers Union warned of “cognitive surrender, ” arguing algorithmic oversight can corrode skills and autonomy. John Chadfield, national tech officer at the CWU, said the public would be “horrified to hear of bosses ‘deploying deeply intrusive surveillance on workers’” and called for laws to ensure AI is used “in a socially responsible way and not to ‘shut off job opportunities for young people, degrade the craft of skilled workers and make astronomical profit for a few’.”

Where the rollout tends to go wrong, technical and governance gaps

The technical pipeline is simple: automatic speech recognition turns audio into text, an LLM or classifier tags behaviour in the transcript, and a dashboard shows scores and notes to managers. That chain works today, but it has predictable failure modes that matter in a workplace.

Security researcher Simon Willison has argued that as models became broadly useful in 2025, the real risks shifted from raw capability to deployment architecture and governance. He highlights threats such as prompt injection and a dangerous combination he calls the “lethal trifecta”: systems that (a) access private data, (b) accept untrusted external instructions, and (c) can exfiltrate information. Willison also warns of normalisation of deviance, where organisations grow complacent as risky architectures keep working until they fail.

Design patterns that reduce risk exist. Willison references DeepMind’s CAMEL architecture and similar approaches that isolate sensitive functions into quarantined components, apply taint-tracking and require explicit human approval for consequential outputs. These are practical mitigations employers should be able to describe when they deploy monitoring agents.

Operationally, simple errors cause big harms: ASR mistakes such as misrecognised words and accent bias, miscounted filler phrases, rubric rules that punish visible self-correction, and miscalibrated confidence scores. Any of these can create false positives that push managers toward unwarranted interventions, and they erode trust fast.

What responsible deployment looks like

Quality assurance and scalable feedback are legitimate goals. A responsible rollout must show the tradeoffs and put guards in place. Practical elements include:

  • transparency about what data is recorded, who can access it and whether recordings/transcripts may be used to train models;
  • early access for instructors to their own session data, the AI’s notes and a clear, timely appeals process;
  • published performance metrics for the system, confusion matrices and false positive/negative rates measured in a pilot against multiple human raters;
  • strict retention policies, role-based access controls and explicit vendor commitments about not using raw session data for model training unless consented and contractually protected;
  • human-in-the-loop gates for any consequential HR action, with audit logs showing when managers followed AI recommendations and why;
  • pilot periods run with union and staff consultation, plus wellbeing support and monitoring for mental-health impacts during rollout.

These are not theoretical safeguards. They map directly to the practical failure modes that produced the teacher complaints described above.

Questions you should be asking, and candid answers

  • Is Multiverse using AI to evaluate instructors?

    Yes. The system records, transcribes and scores sessions, assigns a “risk status” and produces written comments for managers, according to reporting and company statements.

  • What behaviours does the AI flag?

    Reportedly: slow handling of connection problems (benchmark ~1 minute), repetitive filler phrases or hedges, vague answers, learners going quiet and prolonged one-to-one sidetracking. The rubric that staff described emphasizes judging patterns across sessions rather than isolated instances.

  • Are managers still the final arbiter of performance?

    Multiverse says human managers write performance reviews and that the AI “directs human time” rather than replacing judgment. Teachers, however, report managers acting on AI flags while staff lack early access to the dashboard, a factual dispute that requires audit logs to resolve.

  • Is this causing harm to staff?

    Multiple anonymous instructors report increased stress, sleep disruption and a sense of constant surveillance. The union has warned of broader risks such as “cognitive surrender.” These are credible reports of harm but not clinical causal findings; employers should treat them as urgent signals to investigate and mitigate.

  • What technical risks should employers be worried about?

    ASR errors and accent bias, miscounted filler detections, rubric misinterpretation (e.g., penalising self-correction), miscalibrated confidence scores, prompt injection or input manipulation, and data-exfiltration risks when systems touch sensitive recordings. Simon Willison’s analysis and the quarantine/taint-tracking patterns he cites are a useful technical reference.

  • What remains unclear or unverified from public reporting?

    Key unknowns include the system’s measured accuracy (false positive/negative rates), whether recorded sessions are retained or used for model training, the precise vendor contracts and data flows, and whether AI flags have already driven disciplinary HR outcomes. These require documents or audits from the company and regulators.

Three practical next steps for leaders

  • Pause high-stakes use until audited. If the system feeds into performance management, pause any consequential HR actions until an independent audit produces calibrated error rates and a human-in-the-loop policy.
  • Publish pilot metrics and open access. Release confusion matrices, FPR/FNR by rubric category for the pilot period, and give instructors immediate access to their session data and the AI’s annotations.
  • Contractual and wellbeing protections. Commit in writing that raw session data will not be used to train external models without consent, limit retention, implement strict access controls, and provide support (wellbeing resources and an appeals path) during rollouts.

Why this matters beyond one company

Automated scoring changes incentives. Young professionals can learn to perform for the algorithm rather than adopt best professional practice. Managers nudged by numerical risk labels can default to punitive responses. Logged sessions reused for model training risk creating a closed feedback loop that flattens craft into a statistical caricature.

If organisations want to scale feedback with AI, they must also scale transparency, auditability and worker protections. Otherwise the immediate cost will not be a single misgraded lesson, it will be burned-out staff, eroded trust and a degraded profession.