Tech Landscape #439
Four new image models, Facebook’s TikTok plans, and the rise of voice control
Hello!
This week Substack announced that it has added an AI text detector, estimating how much of any post’s content was human-written. Because, it says:
The core problem is not people using AI, or the quality of its output. Not everything made with AI is slop, and not all slop is made with AI. The problem is when there is a mismatch between a reader’s expectation and reality, especially when they unwittingly invest their attention in something with no human thought on the other end.
Largely, I think this is fine; I’ve been transparent about using AI to assist me with putting together this newsletter (I wrote about my full workflow in issue 435) and I don’t want to read the flat prose output by low-effort prompts any more than you do. My worry is the reliability of AI text detectors; in the febrile atmosphere of social platforms, a false-positive result could easily lead to a witch hunt. And the reason for my worry is because I’ve got proof of a false-negative already; this is the result I got when I tested this issue with their tool:
I’d estimate that ~20% of the writing below is generated. Perhaps it didn’t flag because I use a prompt with examples of my own writing to try to match my style; if that’s the case, then the detection tool is very easy to fool.
Anyway, let’s get on with it. Hope you’re well!
Synthetic Content
Black Forest Labs announced FLUX 3, a unified multimodal model that can generate video (up to 20-seconds with native audio), image, and action prediction for physical AI / robotics. bfl.ai
It’s currently in invite-only early access, so I’ll write more about it when it’s released more widely.
Midjourney released v8.2, focused “on aesthetics and image quality and personalisation”. midjourney.com
I haven’t had the chance to give this a proper run-out (please stop doing Friday releases, tech companies), but I knocked up a quick side-by-side example ⬇️ showing 8.1 vs 8.2 vs 8.2 + personalisation with the same prompt and seed. I would note that Midjourney’s focus on aesthetics makes it remain the model of choice for many artists, despite there being many technically ‘better’ models on the market now.
Microsoft introduced MAI-Image-2.5-Pro, a highest-quality variant of its image model with editing and precise in-image text rendering. It also released the previously-announced MAI-Voice-2-Flash, a faster, lower-cost voice model. Both are in public preview. microsoft.ai
I made this ⬇️ example; it’s pretty decent, nice level of detail, but the real-world grounding is a bit suspect (Big Ben has swapped sides). MAI-Image has replaced GPT-Image across Microsoft’s major creative surfaces, including Bing Create and Powerpoint; with this and the news that MAI-Code-Flash is now in Excel and GitHub Copilot, its reliance on OpenAI is coming to an end.
Qwen launched Qwen-Image-3.0, designed for more detailed images, with complex visual layouts and precise small-text rendering in 12 languages. qwen.ai
Eachlabs launched Flint Image, an image model that generates four distinct visual variations from a single prompt for roughly the price of one, using Gemini for reasoning before each generation. eachlabs.ai
Audio & Avatars
ElevenMusic introduced Vocals and Styles, letting users generate songs with a consistent voice and audio style, using their own recordings or pre-made library options. elevenlabs.io
It also released the Finetunes Music API, allowing developers and brands to train custom music models on their own audio tracks for consistent brand and musical identities. elevenlabs.io
Seed Audio now allows time stamp tagging, enabling detailed time control over generated speech and audio. x.com/BytePlusGlobal
Synthesia released Dubbing 2.0, an updated video localisation stack that brings better lip sync, more natural voices, smarter translations, and a faster editing workflow. synthesia.io
Tavus introduced Presentation Mode, enabling its AI avatars to present slide decks live during interactive calls, with options to walk through materials sequentially or surface specific slides on demand. tavus.io
Creative Tools
Reve introduced Templates, letting users swap products, images, or text while maintaining layout consistency, automatically updating surrounding elements to keep the image cohesive. blog.reve.com
This company (which I learned is pronounced rev, not reeve) is really on a tear at the moment. It also added new video models (Kling 3 and Seedance 2 / Fast), and released a nice promo video.Runway added Workflows in Agent, letting users build, run or edit node-based workflows through natural language. instagram.com/runwayapp
Runway launched Media Router, a tool in Runway Dev that automatically selects the best video, image, or audio model for generation requests, optimising options based on user preferences for cost, quality, and latency. runway.com
This kind of automated model switching depending on the task is quite common in assistants and LLMs, less so in creative tools.Flick added 3D Stage, using a 3D scene layout for composition to be rendered with an AI image model. instagram.com/flickartai
It’s a kind of built-in mini-Blender, with AI as the rendering pass. This technique is starting to gain momentum.
Social
Meta Apps
Some interesting moves from Facebook: first, users can apply to be Verified, which gives a (free) badge that shows proof of identity and good standing — especially useful for dating and selling.
Next there’s the release of Seller, a purpose-built app for Marketplace sellers, with a unified inbox, inventory management, performance insights, and AI-powered listing creation. It follows on the heels of other purpose-built apps: Forum [TL 430], which is fully synced with Groups plus adds new management tools; and Creator Studio [TL 435].
And sort of buried in the announcement is the news that Facebook will test a full-screen video-first experience — a lá TikTok — in select markets where video consumption is high. This may be part of a long-term strategy that explains why core features are being spun off into separate apps.
Threads now lets users reorder carousels before posting, in the mobile apps. threads.com/@threads
Instagram now lets users swap audio in older posts without having to delete and reupload. instagram.com/creators
These are ‘papercut’ fixes; small, but useful.
YouTube
YouTube introduced new thumbnail creation features, letting Partner Program creators upload custom thumbnails for Shorts, and everyone generate thumbnails for long-form videos using Ask Studio. blog.youtube
YouTube updated Communities with Community-only posts and more moderation options for creators, and quote posts and more discovery surfaces for fans. youtube.com
Safety
Threads added parental supervision tools, letting parents track screen time, set daily usage limits, adjust sleep mode, and manage privacy settings, via Meta’s Family Center. about.fb.com
Twitch introduced optional parental controls for teen accounts; parents can restrict streaming, disable direct messages, set daily time limits, and filter content recommendations. safety.twitch.tv
If these features had been in place a long time ago, perhaps social platforms wouldn’t have needed the heavy hand of regulation.
Assistants & LLMs
Google launched three new Gemini models: 3.6 Flash, the new standard ‘workhorse’ model with improvements to coding, knowledge work, and multimodal performance; 3.5 Flash-Lite, the fastest and most cost-efficient variant; and 3.5 Flash Cyber, a specialised security model. blog.google
Improved token efficiency and faster output speeds are very much at the forefront here, because of Google’s scale. The company is taking flak in some quarters because it doesn’t have a model that competes at the very top level, such as Claude 5 Fable and ChatGPT 5.6 Sol. But it has very different considerations; it has to support billions of people, vastly more than its rivals. It needs to be good enough for most people, not the very best for some. That said, Gemini 3.5 Pro is in testing, and Gemini 4 is in training, so it will be interesting to see if it tries to reach that top tier.
Gemini Spark is coming to Pro subscribers (in the US) and Gemini Intelligence now works with more apps (in the US and Korea).
Anthropic launched Claude Opus 5, which it says performs near the level of Fable 5 — and even exceed it at coding, knowledge work, and visual outputs — at half the cost. anthropic.com
Meta AI gained new agentic capabilities that make it more of a personal assistant, enabling it to plan and research, connect to email and calendar apps, create slide presentations, and more. about.fb.com
Grok 4.5 is more broadly available across the Web, X, iOS, and Android. x.ai
Agentic Work
ChatGPT brought the new Voice mode to its desktop app for controlling agentic Codex and Work sessions through speech. instagram.com/openaidevs
Claude’s voice mode now works with its latest models, and users can teach Skills to Claude Cowork by screen-recording a task while narrating it.
I’m slowly coming around to using voice dictation / instruction more. It feels a bit weird; talking is such a natural mode of communication, but it takes some cognitive effort to start doing it with computers because the keyboard has been dominant for so long.
Manus introduced Plan Mode, an opt-in feature that generates an editable, structured plan for user review and approval before it executes complex prompts. manus.im
Genspark launched AI Workspace 6.0, bringing many of its work tools together in a unified platform featuring the new Second Brain persistent memory layer that retains context across tools, plus new design and email creation tools and collaborative human-agent workspaces. genspark.ai




