Procurement must mandate AI disclosure after flawed citations in A$3.48M ACCS report

Dubious footnotes undercut a A$3.48 million trial, and procurement needs to catch up

Public policy needs reliable evidence. A A$3.48 million age‑assurance technology trial run by the UK‑based Age Check Certification Scheme (ACCS), a report that helped inform Australia’s under‑16 social media ban, is now under scrutiny after investigators flagged multiple problematic references and apparent signs of generative‑AI involvement.

What happened

Guardian Australia analysed the ACCS trial report and, in a submission to a Senate inquiry, found citation problems. They range from DOIs that resolve to the wrong paper or nothing at all, to author/journal/year combinations that do not match any record. Guardian reporting found at least six problematic references in one section and said four links in Parts E and K (sections of the ACCS report) contained metadata traces investigators say suggest ChatGPT was used to generate or rewrite content.

ACCS initially denied using AI in the generation of the report. An ACCS spokesperson later said, “We did not use AI in the generation of the report or the cited materials … Each one of them were checked as genuine links and reports and that they were relevant to the specific issue being cited.” After Guardian investigators reported metadata traces, ACCS acknowledged that ChatGPT had been used to “rewrite paragraphs more succinctly, ” while keeping that links were human‑verified. In a later statement the ACCS spokesperson said: “All of the links were checked and verified and, whilst I apologise if there is an inadvertent error, this was all done by human verification.”

The department responsible for the contract told senators that ACCS “assured us that the links were all checked at the time that the document was published, and they all worked then, ” adding, “So, something may have changed in the interim. [They] did also assure us that while some of the links may have been broken, though the source documents do still exist, they are still valid.” The communications minister, Anika Wells, described the ACCS report as showing “many effective options” for age checking and said it “paved the way” for the under‑16 ban when the law came into effect in December last year.

Exactly what kinds of errors were reported

  • DOIs that point to non‑existent articles or to different papers than the citation suggests.
  • Author/journal/year combinations that do not match publisher records.
  • Citations used to support claims where the cited paper does not in fact back the assertion.
  • An instance where a paper was cited as “accessed” in March 2025 despite its lead author saying it was not publicly available until June 2025; the same citation listed a person named “Jamil” who was not on the published paper, according to reporting.
  • Four hyperlinks in Parts E and K for which Guardian investigators reported metadata traces they say point to ChatGPT as the source of generated or rewritten link content; the technical specifics of that forensic claim have not been publicly detailed.

What is confirmed and what still needs independent verification

  • Confirmed by reporting: the trial cost was A$3.48 million. Guardian Australia identified multiple citation problems and reported metadata traces linking some links to ChatGPT. ACCS first denied then acknowledged limited ChatGPT use. Senators and experts have raised concerns.
  • Not yet independently confirmed in public material: the complete scope of incorrect citations across the full 1, 000‑page report. Whether each erroneous citation was generated by AI versus introduced through human error or link‑rot. Whether any specific flawed reference materially altered the report’s recommendations that influenced policy.

Why these errors matter

At a surface level, bad references are sloppy. At a policy level, fabricated or misleading citations can erode trust, create the impression of peer support for unsupported claims, and spread falsehoods into the policy ecosystem. Independent voices made the point bluntly: Senator Fatima Payman criticised what she called the “saga of last year’s Deloitte AI slop report” and demanded stronger accountability from contractors, while Prof Christian Downie (ANU) warned that procurement must include penalties and verification rights to discourage this behaviour.

There is a recent, directly relevant precedent. According to The Guardian (6 Oct 2025), Deloitte refunded part of a A$440, 000 government contract after admitting generative AI had been used in a flawed government report that contained nonexistent references. That case shows procurement can include material consequences when AI‑assisted deliverables contain demonstrable errors, though each case depends on contract terms and the specifics of the work.

What boards, procurement teams and counsel should do right now

Treat this as a solvable governance problem rather than a moral panic about tools. Prioritise immediate, enforceable safeguards:

  1. Require AI‑use disclosure up front. Clause example: “Contractor must disclose all generative AI models used in producing deliverables, provide model prompts/outputs that informed any text, and certify the exact locations where AI outputs were incorporated.”
  2. Make citation verification mandatory before acceptance. Clause example: “All citations must resolve via CrossRef/doi.org to the cited work, match publisher metadata (authors, year, title), and be supported by a saved copy of the source or a publisher’s landing page screenshot.” Why this matters: CrossRef and publisher records are the canonical way to validate DOIs and publication dates.
  3. Preserve an immutable audit trail. Require deliverables to include original source files, timestamps, and a version history showing edits. Clause example: “Provide working files, editor logs, and an export of document metadata at submission; grant limited audit rights to verify provenance.”
  4. Define contractual remedies and enforcement. Practical remedies include financial clawbacks for demonstrable falsehoods, mandatory corrections and resubmissions, and the right to withhold final payment pending an independent forensic audit. The Deloitte reporting shows refunds are a real outcome when errors are material and provable.
  5. Train and appoint forensic‑capable reviewers. Don’t rely solely on general editorial staff. Assign reviewers trained to run CrossRef/DOI checks, inspect PDF metadata (exiftool or equivalent), and contact corresponding authors when publication dates or authorship appear inconsistent.

Quick, actionable checklist for three audiences

  • Procurement teams: amend active contracts to require AI disclosure and citation verification; add acceptance gates that include CrossRef resolution and saved source artifacts.
  • Boards and executives: insist on documented audit trails for critical policy work and a named accountable officer for evidence verification before sign‑off.
  • Journalists and parliamentarians: request working files, editor logs and the report’s PDF metadata when investigating contested references; demand independent forensic audits when errors could materially affect policy.

Short Q&A: common questions and honest answers

  • Were AI tools used in the ACCS report?

    Guardian Australia reported metadata traces in four links that investigators say suggest ChatGPT was used to generate or rewrite content, and ACCS later acknowledged using ChatGPT to “rewrite paragraphs more succinctly.”

  • What kinds of citation errors were identified?

    Investigators reported DOIs pointing to the wrong or no paper, mismatched author/journal/year metadata, and citations that do not support the claims attributed to them; Guardian found at least six problematic references in one section.

  • Did the citation errors change the policy outcome?

    That is not clear. The department says the report “paved the way” for the under‑16 ban and that ACCS assured links worked at publication. Whether any flawed citation materially altered recommendations requires a detailed audit of the report and its decision‑making chain.

  • Is there a procurement precedent for penalties?

    Yes: according to The Guardian (6 Oct 2025), Deloitte refunded part of a A$440, 000 contract after admitting generative AI had produced errors in a government report. That case shows financial remedies are possible, though outcomes depend on contract terms and the severity of errors.

  • What immediate steps should organisations take?

    Audit recent contractor reports for provenance, require AI‑use disclosure in active contracts, and add mandatory citation validation and preserved working files to acceptance criteria.

Last word

Generative models are powerful drafting tools, but they are not sources of record. The way forward is simple: build provenance and verification into procurement and governance. Do that, and you keep the productivity upside of AI while protecting the credibility that good policy depends on.