Executive summary: Generative images are fast, but when every burger looks identically glossy, customers notice and conversions fall.
Scroll a delivery app and you’ll see it: plates that are too glossy, garnishes frozen in identical poses, textures that resolve into an unsettling sameness. TechCrunch reported on 2026/09/03 that this pattern is emerging across AI‑generated restaurant imagery. The reasons are technical (how models are trained), practical (how teams use and edit outputs), and commercial (customers trust what looks convincingly real, and they react when it doesn’t).
How those images actually get made
Most food images produced at scale today come from two linked components: diffusion‑based image generators (the pixel engines behind tools like Midjourney and similar systems) and language models that craft prompts, captions, or multimodal instructions. Diffusion models build images by progressively denoising a noise field guided by text. LLMs typically provide the prompts, structure menu copy, or drive multimodal pipelines.
Two practical consequences follow. First, models learn and reproduce dominant patterns in their training data. Second, human operators often run many prompt and edit cycles and pick the “best” outputs. That process favors the safest, most polished images and reduces variety.
Why so much “sameness” shows up
- Training‑data bias: If a model’s corpus is rich in commercial, well‑lit food photography, think chain‑restaurant shots, the model will prefer those signals. It learns what’s typical and samples toward that mode.
- Optimization for pleasingness: Objective functions and sampling choices often favor high‑fidelity, inoffensive outputs (lower diversity, higher likelihood). That “shaving off the edges” produces smooth, non‑challenging images rather than idiosyncratic ones.
- Iterative smoothing by humans: Repeated prompting, retouching, and A/B selection push outputs toward a consensus aesthetic. Over many cycles, quirks get edited out and artifacts that look pleasing to reviewers get amplified.
As Lee Rainie observed to TechCrunch, “The optimization of the data sets is for pleasingness, or you know, not being offensive, and so there’s a way that turns into homogenization… What AI is known to do both in images and language is to shave off the edges.”
Colorful metaphors, and what they really signal
TechCrunch relayed several memorable lines from Alex Lisle, CTO of Reality Defender, that capture business risks without being literal technical diagnoses. Lisle said, “It’s almost like an alien trying to make a pizza without understanding its core principles, ” and added, “A lot of this stuff looks like a Chili’s menu from 2015, and there’s a reason for that.” He also warned, metaphorically, that “model collapse is almost like a mad cow disease… when you feed the outputs from one model back into itself, eventually the inbreeding becomes too much, and the whole thing collapses, ” and he noted that what we’re seeing today is convergence rather than collapse.
Those metaphors are useful shorthand. Convergence, the visible loss of stylistic diversity, is observable now. Full‑scale “model collapse, ” where systems materially degrade because their training pools become dominated by synthetic outputs, is a debated, longer‑term risk that depends on data‑collection and curation practices across the industry.
Why the images feel wrong to people
It isn’t just subjective taste. Researchers at the University of Duisburg‑Essen reported an “uncanny valley” effect for near‑realistic AI food images: pictures that are very close to plausible but wrong in subtle texture or lighting details can provoke stronger aversion than obviously fake images. Social reactions mirror that finding: an X user named Labtec edited a ChatGPT‑generated menu “100 times” and wrote, “The end result actually makes me uncomfortable.”
Perception matters because customers equate realistic imagery with reliability. As Alex Lisle put it, “Seeing and hearing has always been believing… That’s no longer the case. The world has fundamentally shifted, for good or for ill.” When diners suspect imagery is synthetic and misleading, trust and orders can evaporate quickly.
Market responses and provenance pressure
Startups offering detection and provenance tools (Reality Defender among them) are racing to meet demand. Platforms and vendors have split responses. Some leaned in to cut costs and scale catalog imagery. Others pulled back after customer complaints. Reporting tied to this topic has also noted aggressive data‑collection practices. One report said companies have at times scanned and ingested rare or proprietary material into training pools (a high‑profile example was reported on 2026/08/17). This underscores why provenance and audit trails matter.
Practical steps for restaurant and brand leaders
AI can speed production and lower costs, but value comes from the combination of machine scale and human judgment. Here’s a pragmatic playbook that moves management from reactive to deliberate.
- Keep real photography where it counts: Reserve professional food photography for hero menu items, high‑traffic pages, and paid placements. Use AI for ideation, conceptual mockups, or internal menus where the risk is low.
- Mandate human QC and clear KPIs: Require a creative director or food stylist to approve customer‑facing images. Track CTR, add‑to‑cart rate, order conversion, dwell time, and complaint/return rates on items using AI images versus real photos.
- Run a short A/B test before rollout: Randomize real photos vs AI images for the same items for a defined period. Treat the real‑photo baseline as the success threshold. Require non‑inferiority, or a statistically significant lift, on conversion metrics before swapping images at scale.
- Demand provenance and content credentials: Insist vendors attach Content Credentials (C2PA or similar) and metadata to synthetic assets. Embed provenance in files and contracts so you can trace sources and exclude synthetic outputs from public training pools.
- Archive and label synthetic files: Keep AI outputs in a segregated, labeled dataset and forbid vendors from contributing them to public corpora. That reduces the feedback risk of contaminating future model training.
- Use detection tools plus manual review: Automated detectors help at scale but are an arms race. Combine them with spot checks and consumer panels to catch subtle “off” imagery.
- Contractual protections: Require vendors to disclose training‑data provenance and to prohibit reuse of your images in public training sets. Add audit rights for high‑impact suppliers.
A simple A/B test template for busy execs
- Duration: 1-2 weeks, depending on traffic volume.
- Randomization: Randomly assign users to “real photos” or “AI images” groups for identical menu items.
- Metrics: Measure CTR to item page, add‑to‑cart rate, conversion to purchase, average order value, dwell time, and complaint rate.
- Decision rule: If AI images are statistically non‑inferior to real photos on conversion (standard threshold p<0.05 or equivalent business rule), consider limited rollout; otherwise, revert and refine.
Three quick operational guardrails
- Always require human sign‑off for customer‑facing images.
- Attach content credentials (C2PA) and retain provenance metadata.
- Archive synthetic content separately and block it from public training pools by contract.
Key questions, short, honest answers
-
Why do AI‑generated menu photos all look similar?
Models trained on large collections of commercial photos learn the dominant, “safe” aesthetic; optimization choices and repeated human selection push outputs even closer to that mode, producing homogenized images.
-
Is this only an aesthetic problem?
No. It’s aesthetic, commercial, and technical: customers can distrust images (hurting conversion), brands lose distinctiveness, and unchecked recycling of synthetic media risks longer‑term dataset contamination and reduced model diversity.
-
How worried should we be about “model collapse”?
Model collapse is a debated, worst‑case outcome when synthetic outputs dominate training pools. What’s observable now is convergence (loss of variety). Mitigations, provenance, filtering, and vendor contracts, make collapse avoidable in practice.
-
Can detection tools solve the problem?
Detection and provenance tools reduce risk but aren’t foolproof. Combine automated checks with manual QC, provenance metadata (C2PA), and contractual restrictions to make a durable defense.
-
Should restaurants stop using AI for menus?
Not necessarily. Use AI strategically for ideation, low‑risk assets, and scale, but keep customer‑facing, revenue‑critical imagery grounded in human photography and rigorous testing.
Visuals are a brand’s handshake. AI will speed that handshake, but only thoughtful processes keep it feeling human. For leaders, the priority is simple: measure, mandate provenance, and never let automated scale replace human taste. When every burger looks like the same glossy double, the first casualty is not art, it’s conversion.