# Creative & Generative Media

AI for generating and editing images, video, audio, 3D assets, and cross-media content. Mostly leading-edge with rapid advancement — image generation, music composition, and voice synthesis are approaching good practice. Video generation and 3D asset creation are progressing fast but quality and controllability gaps persist. The most active domain by momentum: over half the practices are advancing.

> Creative AI is now a feature inside software you already pay for, and the models underneath can vanish. Tests published this fortnight show fully generated ads and dubs still losing to human-finished work.

## The Picture

The capability question is settled: in blind tests listeners cannot reliably tell generated music from human, and cloned voices fool most people. The acceptance question is not. Most marketing teams now use AI for backgrounds, resizes, variants and training video, with a person approving what ships, and if that describes you, you are in the pack. The better results on record come from teams that train models on their own products and keep human finishing; the exposed ones publish fully generated work to audiences who keep telling surveys they trust it less. Vendors prosper either way: Adobe reports annual recurring revenue from AI-first products above $650 million.

## This Fortnight

- **Adobe closed its roughly $340 million purchase of Topaz Labs on September 23.** Topaz's image-enlarging tool now runs inside Photoshop and Lightroom against a metered credit balance, and Adobe has said nothing about customers who bought Topaz outright under perpetual licenses. Specialist tools are becoming line items in subscriptions you already hold, so ask your creative teams which standalone licenses they depend on and what the credit meter will cost at their volume.

- **OpenAI cut off developer access to its Sora 2 video model on September 24 and named no replacement.** The consumer app had already closed in April after earning $2.1 million in total. Any workflow built on a single video model needs a tested fallback, because a vendor can retire a model without offering anything to move to.

- **ElevenLabs released a model on September 28 that clones a voice from 10 seconds of audio.** A share sale in the same fortnight doubled the company's valuation. Cheaper, faster cloning helps localization and customer service, and it equally helps anyone impersonating your chief financial officer, so payment approvals that rest on recognizing a voice need a second check.

- **Incode reported that a leading public deepfake detector fell from 91.2 percent on academic benchmarks to just over 60 percent on its own identity-verification data.** Incode, which sells identity verification, attributes the gap to the difference between polished public test images and unfiltered real-world selfies. Treat detection software as one fraud signal among several, never as a verdict on whether a face or voice is real.

- **In one paid-media test, 20 fully AI-generated ad variants finished 40 percent below two hand-finished ads on click-through.** The winners paired a model trained on the brand's own products with human copy, retouching and layout, and a dubbing vendor with an interest in the answer reported YouTube's automatic dubs holding viewers far less than professional ones. Budget for human finishing on anything customer-facing, and judge these tools on results, not volume.

## Coming Up

- **Google starts retiring an older Gemini image model on October 2, and Photoshop has already begun removing it.** Adobe's editing features now route work across its own and partner models, so a retirement upstream can change what your team's saved workflows produce. Have someone list which models your recurring creative jobs depend on, and re-test output whenever one changes.

- **EU rules requiring disclosure of AI-generated media have applied since August 2, and the duty sits with the advertiser, not just the agency.** Fines run up to €15 million, and European Commission guidelines treat hidden metadata alone as insufficient. Check before your next European campaign who in your agency contracts is responsible for visible labeling, and watch the first enforcement cases for how strictly regulators read the rule.

- **YouTube plans to pilot real-time dubbing for livestreams in early 2027.** Languages and technology are undisclosed, its existing automatic dubbing is already on by default, and performers' consent terms are still being fought over in courts and union contracts. Before relying on default dubs in a new market, compare viewer retention by language track and keep professional dubbing for humor, children's and emotional content.

## What's Hard About This

- **Generation is cheap, but a publishable asset is not.** One industry assessment found 68 percent of enterprise video deployments still need human intervention to reach publishable quality, and review capacity has not grown with output. The saving is real only if you count the checking, so measure cost per accepted asset, not cost per generation.

- **Your audience judges the output, and much of it is wary.** Gartner finds half of consumers would rather buy from brands that keep AI out of public-facing content, and platforms are cutting reach and payouts for generated material. No tool upgrade fixes that; it is a decision about where in your customer experience synthetic content is acceptable and how it is disclosed.

- **Marking synthetic media is now a legal duty, while detecting it keeps failing against current generators.** Detectors that reach 94 percent on older test footage manage 48 to 70 percent on video from today's tools, and labels are stripped as content moves between platforms. A missing mark does not prove something is human-made, so verifying anything that matters, from a supplier's video call to a news clip, has to rest on process, not software.

## Practices (21)

- [3D asset, scene & texture generation](https://www.thestateofplay.ai/practice/3d-asset-scene-and-texture-generation) — Leading Edge, Steady
- [AI-driven video editing & post-production](https://www.thestateofplay.ai/practice/ai-driven-video-editing-and-post-production) — Leading Edge, Steady
- [Audio production — editing, podcasts & sound design](https://www.thestateofplay.ai/practice/audio-production-editing-podcasts-and-sound-design) — Leading Edge, Steady
- [Avatar generation & personalised media at scale](https://www.thestateofplay.ai/practice/avatar-generation-and-personalised-media-at-scale) — Leading Edge, Steady
- [Brand asset generation & variation](https://www.thestateofplay.ai/practice/brand-asset-generation-and-variation) — Good Practice, Steady
- [Content authenticity — deepfake detection & provenance](https://www.thestateofplay.ai/practice/content-authenticity-deepfake-detection-and-provenance) — Leading Edge, Steady
- [Image editing — inpainting, outpainting & extension](https://www.thestateofplay.ai/practice/image-editing-inpainting-outpainting-and-extension) — Leading Edge, Steady
- [Image editing — style transfer & artistic transformation](https://www.thestateofplay.ai/practice/image-editing-style-transfer-and-artistic-transformation) — Leading Edge, Steady
- [Image generation — photorealistic & illustrative](https://www.thestateofplay.ai/practice/image-generation-photorealistic-and-illustrative) — Good Practice, Steady
- [Image generation — product visualisation & mockups](https://www.thestateofplay.ai/practice/image-generation-product-visualisation-and-mockups) — Good Practice, Steady
- [Image processing — upscaling, restoration & compositing](https://www.thestateofplay.ai/practice/image-processing-upscaling-restoration-and-compositing) — Leading Edge, Steady
- [Interactive content — game, AR/VR & environment generation](https://www.thestateofplay.ai/practice/interactive-content-game-arvr-and-environment-generation) — Leading Edge, Slowing
- [Lip sync & video dubbing](https://www.thestateofplay.ai/practice/lip-sync-and-video-dubbing) — Leading Edge, Steady
- [Motion capture & pose estimation for content production](https://www.thestateofplay.ai/practice/motion-capture-and-pose-estimation-for-content-production) — Leading Edge, Steady
- [Multimodal content generation](https://www.thestateofplay.ai/practice/multimodal-content-generation) — Good Practice, Accelerating
- [Music generation — background & ambient](https://www.thestateofplay.ai/practice/music-generation-background-and-ambient) — Leading Edge, Steady
- [Music generation — full composition](https://www.thestateofplay.ai/practice/music-generation-full-composition) — Bleeding Edge, Steady
- [Text-to-speech — natural voice synthesis](https://www.thestateofplay.ai/practice/text-to-speech-natural-voice-synthesis) — Leading Edge, Steady
- [Text-to-speech — voice cloning & custom voices](https://www.thestateofplay.ai/practice/text-to-speech-voice-cloning-and-custom-voices) — Leading Edge, Steady
- [Video generation — long-form narrative & explainer](https://www.thestateofplay.ai/practice/video-generation-long-form-narrative-and-explainer) — Bleeding Edge, Steady
- [Video generation — short-form](https://www.thestateofplay.ai/practice/video-generation-short-form) — Leading Edge, Steady

## Full Technical Briefing

## Where AI Stands in Creative & Generative Media

Creative and generative media is the domain where the capability question has been answered most decisively and the acceptance question least. In a blind test published in the NeurIPS Creative AI track, listeners could not reliably tell Suno output from human music. Voice clones reach 97% fidelity while humans detect only 37.5% of them. The leading text-to-video models sit in a statistical tie on independent arenas. Money has followed: Adobe reports AI-first annual recurring revenue above $650M, up 150% year on year; ElevenLabs says it is pacing at $600M and was valued at $22bn in a secondary share sale; Runway is at $200M, Kling reports $300M annualised and Suno $300M. Yet the same period produced a casualty list. OpenAI closed Sora, an app that earned $2.1M in total. Beatoven.ai shut its consumer music service in August. Soul Machines entered receivership in February. Adobe's share price fell from $630 in February 2024 to $266 by September 2026 even as its AI revenue grew.

What works in production is bounded. WSC Sports generated more than 134,000 videos across the 104 matches of the 2026 World Cup; transcription and noise removal are dependable on clean audio; avatar video is replacing filming for corporate training and localisation; brands generate backgrounds, resizes and variants around a real photograph. Wherever output must be faithful to something, such as a product, a face, a brand colour, a storyline or a mesh that a rigger can use, a person still stands between generation and publication. Algorithmine finds 68% of enterprise video deployments need human intervention to reach publishable quality. Retailers manually correct 60-80% of inpainting outputs. Photoroom's 850-product benchmark found frontier editing models preserve product accuracy in 29% of outputs. Two areas sit outside that pattern. Workflows that produce text, image, audio and video from one brief are spreading fastest, because Adobe and Google have put them behind conversational assistants and into Google Ads and Workspace, although the UK's ONS still puts visual content creation at 16% of businesses with ten or more employees. Full-song generation is the reverse case: enormous consumer volume, almost no organisations reporting that it works in production, and 87% of Canadian musicians in a University of Alberta study viewing generative AI negatively.

Momentum is building in distribution, not in models. Capability now arrives as a feature inside software people already pay for: five video models in Premiere's timeline, Lyria in Gemini, auto dubbing switched on by default at YouTube, markerless capture in Unreal Engine 5.8. That widens reach and squeezes specialists. The stall is on the demand side. Gartner finds half of consumers would rather buy from brands that keep AI out of public-facing content. Deezer says fully AI-generated tracks exceeded half of daily uploads on peak days in July and draw 1-3% of streams. Among GDC respondents, 52% now view generative AI negatively, up from 30% a year earlier, and Roblox has lost daily users for three consecutive quarters while expanding its AI creation tooling. What separates this domain from its neighbours is that its output is judged by audiences, not by the organisation that deployed it, and that it draws on voices, likenesses and catalogues belonging to someone else. Courts, collecting societies and regulators are rewriting the terms while the tools are already in use: EU AI Act transparency obligations have applied since 2 August 2026, and a Munich court ruled against Suno on 31 July.

## What's New, 2026-09-17 to 2026-10-01

The fortnight's clearest pattern was consolidation around platform owners. Adobe closed its roughly $340M purchase of Topaz Labs on 23 September, putting Gigapixel inside Photoshop and Lightroom on metered credits and leaving legacy perpetual licences unaddressed. It extended its creative tools into Google Gemini and expanded them inside Claude, signed Jet2, was named partner to the NHL's 32 clubs with no launch date given, and now lists Veo 3.1, Kling 3.0, Runway Gen-4.5, Luma Ray3.14 and Seedance 2.0 as partner video models, with the two Chinese models limited to individual plans. ElevenLabs shipped Eleven v4 on 28 September, cloning a voice from 10 seconds of audio across more than 90 languages; it ranks first on the Artificial Analysis arena at 1319 Elo against Cartesia Sonic 3.6 at 1276, and a $300M secondary sale doubled the company's valuation to $22bn. In video, the Sora 2 API was cut off on 24 September with OpenAI naming no replacement. Runway launched an ads agent on 30 September that publishes variants to Meta, Google and TikTok and feeds performance data back into the next round. Pika relaunched as a studio that routes requests across other firms' models. Google took its real-time Gemini avatar agent to general availability on 24 September with SynthID on every stream, and the same day published a four-system orchestration layer chaining Gemini and Veo into sequences of up to 10 minutes. That result is scored on Google's own benchmark, is not packaged as a product and has not been independently replicated. Veo 3.1 clips still cap at 8 seconds.

The more instructive evidence was negative, and much of it concerned measurement. Incode reported that GenD, a leading public deepfake detector, fell from 91.2% benchmark AUROC to just over 60% on its own identity-verification data. A self-audited detector ensemble returned a 99.85% probability of AI in a case where four of five detectors were silent and the fifth said "real". Vidmoat found seven bugs in its own video-editing benchmark, including an agent run marked complete whose export was black for 85 of 92 seconds. Krisp's open benchmark of 265 recordings showed voice isolation cutting pooled word error rate by 73% while making clean phone audio slightly worse, from 3.48% to 3.91%. Atlas Cloud, which sells its own endpoints, found only 3 of 36 cloud image-editing endpoints accept a mask file. On outcomes, AIR Media-Tech, a dubbing vendor with an interest in the answer, reported that YouTube's auto-dubbed tracks held viewers 4 to 10 times less than professional dubs across 400+ channels. A practitioner's paid-media test had 20 fully generated ad variants finish 40% below two hand-finished ads on click-through. VML's survey of 28,000 consumers in 17 countries found 49% say AI product images lower their trust in a brand. A peer-reviewed survey of 105 3D practitioners found generated meshes routinely failing rigging, with over a third of those answering spending 2-5 hours per mesh on topology fixes, even as Tripo P2 shipped native quad output and Meshy researchers reported 94.1% of UV seam predictions unwrapping without post-processing. Roblox opened its prompt-to-game tool as a public alpha in three markets, with roughly 9,000 games published, mostly by first-time Studio users, against a third straight quarter of falling daily users. Nothing in the fortnight changed the overall picture: the new evidence sharpened existing limits and did not move them.

## Key Tensions

- **Generation is cheap; a usable asset is not.** Unit prices have collapsed, but the labour has moved downstream. Practitioner analyses put the keeper rate for short video at roughly one in three, a professional composer found under 1% of more than 20,000 generated tracks production-viable, and game studios report motion-capture cleanup and retargeting consuming 30-50% of animation effort. Review capacity has not scaled with output: one survey found 46% of marketing professionals reporting content stuck in review queues.

- **Supply floods channels that audiences and platforms then close.** Volume and attention have decoupled. Qobuz says AI tracks account for 0.38% of its streams and that it demonetises 60% of AI music, while 221,900 new AI dramas were uploaded in China in the first half of 2026 and 1.3% reached profitability thresholds. Platforms are answering with penalties: Tidal stopped paying royalties on fully generated music from 15 July, YouTube bars template-based AI content from Partner Program monetisation, and TikTok's disclosure toggle carries a 30-40% reach penalty.

- **Platform owners absorb the feature and specialists lose pricing power.** Background music, upscaling, dubbing and style transfer now ship inside Premiere, Photoshop, Gemini and YouTube, often against a shared credit balance. Beatoven.ai's founder cited heavily funded entrants when closing its consumer service in August, and Adobe now owns Topaz. Buyers gain convenience and inherit churn: Runway retired two models on 30 July with no grace period, Photoshop began pulling Gemini 2.5 on 10 September ahead of Google's 2 October deprecation, and aggregator routing can change style, latency or cost without the user touching anything.

- **Provenance is mandated while detection keeps failing.** Marking synthetic media is now a legal duty under EU AI Act Article 50 and California's SB 942, and OpenAI, Anthropic and Google attach C2PA credentials and invisible watermarks by default. Detection does not survive contact with current generators: the DF26 benchmark saw detectors fall from 94% AUC on legacy data to 48-70% on modern full-scene video. Credentials are stripped in distribution and a watermark-removal plugin drew 19 thousand stars in a day, so a missing mark does not show that content is human-made, a limit OpenAI's own documentation states for its tools.

- **Rights and consent are being negotiated after the fact.** Licensing is replacing litigation for those large enough to negotiate: Suno's new models are built with licensed music from Warner, BMG and Believe, though nobody has disclosed what artists will be paid. The output itself remains legally thin, since the US Supreme Court declined to hear Thaler v. Perlmutter on 2 March 2026 and purely generated work stays unregistrable in the US, while Adobe's indemnity excludes claims arising from modifying or combining output. Performers are contesting consent directly: nearly 1,000 actors, agents and others signed an open letter over demands that child actors allow AI use of their voices.

## Top 10 Evidence Items

1. **Adobe Owns Topaz Labs as $340 Million Deal Closes: What Changes for Photographers** (news-coverage) — Shows the platform-owner consolidation pattern the briefing calls the fortnight's clearest trend, with legacy licences left stranded as a cost of that absorption. https://www.photographytalk.com/adobe-owns-topaz-labs/
2. **Why deepfake detection benchmarks fall apart on production data — Incode** (case-study) — Demonstrates provenance tooling failing exactly where it is mandated to work, undercutting the EU AI Act transparency regime now in force. https://www.incode.com/blog/deepfake-detection-in-production/
3. **STT handles noise now. It still can’t handle a background voice.** (case-study) — A rare rigorously measured result that complicates the clean capability narrative: the same tool both fixes and degrades audio depending on input. https://krisp.ai/blog/voice-isolation-benchmark/?ref=toolcenter
4. **ElevenLabs’ new v4 speech model supports more expression control and 90 languages** (news-coverage) — Anchors the money-follows-capability side of the tension, pairing voice-cloning fidelity with the valuation and ARR figures the briefing cites. https://techcrunch.com/2026/09/28/elevenlabs-new-v4-speech-model-supports-more-expression-control-and-90-languages/
5. **When Should You Stop Automating Creative and Customer Service?** (opinion) — A concrete paid-media test contradicting the cost-savings case for full automation, supporting the keeper-rate and review-capacity tension. https://www.cmswire.com/customer-experience/when-should-you-stop-automating-creative-and-customer-service/
6. **Is YouTube's AI Dubbing Safe for Your Channel? Auto Dub vs Pro, With Real Retention Data** (case-study) — Direct evidence for the demand-side stall: default-on scaled dubbing actively costs retention, not just quality. https://air.io/en/youtube-hacks/is-youtubes-ai-dubbing-safe-for-your-channel-auto-dub-vs-pro-with-real-retention-data
7. **Beatoven.ai shuts consumer AI music service as core team joins Rusk Media** (news-coverage) — A named casualty that illustrates platform owners absorbing features and squeezing specialists out of the market. https://routenote.com/radar/beatoven-ai-shuts-consumer-ai-music-service-as-core-team-joins-rusk-media/
8. **Inside Roblox’s Massive Bet On AI-assisted Creation** (industry-report) — Captures the split between internal AI tooling scaling and external audience acceptance falling, the core tension of the domain. https://naavik.co/ai-gaming/inside-robloxs-massive-bet-on-ai-assisted-creation/
9. **Topology-based failure modes in AI-generated 3D assets for production rigging pipelines: a practitioner-informed benchmarking framework** (research-paper) — Grounds the 'someone still stands between generation and publication' claim with hard numbers on rigging rework. https://link.springer.com/article/10.1007/s11042-026-21944-w?
10. **Suno made a new AI music model with major labels. Here's what it means** (news-coverage) — Shows rights being renegotiated after the fact, with licensing substituting for litigation while artist pay stays undisclosed. https://www.latimes.com/entertainment-arts/business/story/2026-09-21/ai-music-licensing-deals-labels-artists-who-benefits

_Source: https://www.thestateofplay.ai/domain/creative-generative-media — CC BY 4.0._
