Hello!
Instead of an intro this week I direct you to my blog; I don’t post there often but this week the spirit moved me to write AI & Books: notes from a moral panic.
Also, I’m running a new experiment this week, trialling a generated overview of the top stories. This is my first effort; I’m aware of its defects, but I’ll refine it over the next few weeks, at which point I’ll make a decision on whether or not to keep doing it (if you have any feelings about it you can hit reply and let me know).
Right, let’s get on with it. Hope you’re well!
Synthetic Content
Two big (Chinese) video model releases this week:
Seedance 2.5
Bytedance’s new model can generate clips of up to 30-seconds length (which can be extended twice for 90-second scenes) and 720p resolution, and edit them. It’s fully multimodal, accepting up to 50 reference inputs (30 images, 10 video clips, and 10 audio clips), has timestamp-level editing controls, and is optimised for pro use cases including white-model control and green-screen editing.
seed.bytedance.com/seedance2_5
It’s currently exclusive to Bytedance’s platforms (including CapCut and Dreamina) but coming to other creative tools soon. I made some quick tests here ⬇️; first clip is text-to-video, second is image-to-video, third is audio+image-to-video. Seedance 2.0 was already the best model and this is even better, with an exceptional level of detail.
MiniMax H3
Formerly Hailuo 3, H3 is a video generation and editing model that can create 4-15 second clips in up to 2K resolution with native stereo sound, from multiple references (9 images, 3 video clips, and 3 audio clips).
It’s widely available now. Intriguingly, MiniMax says it will released the weights for developers to run on their own hardware. I ran some quick tests here ⬇️; the results are good, but there are faults (the scale of the man in the first scene, for example.
Of the two, Seedance 2.5 is clearly better — but it’s also a lot more expensive. H3 may not be as good but it’s faster and cheaper, and the company seems to not be positioning it for the very top end of professional production:
H3 is built for advertising, branding, e-commerce, product design, UI/UX, gaming, and more.
Grok Imagine 1.5 gained text-to-video support, image and voice references, and native 1080p output. x.com/grok
Ideogram released P-Image, a family of four models optimised for the best quality-speed-cost tradeoff. x.com/ideogram_ai
Good enough quality, decent speed, low cost; these models will be a very attractive option for many tasks.
Google Earth integrated image generation, letting Web users generate custom imagery directly on real-world satellite and 3D maps. blog.google
You can restyle images, as I’ve done here ⬇️, or imagine what they’d look like under different circumstances, or add objects in… I can’t imagine using this much, but perhaps it’s not for me. As you can see in my example, usage is limited to three images per day for non-paying users.
Or rather… it was. A day after release, Google pulled the feature to “work on implementing stronger guardrails”, because some people found out that you could manipulate images to, for example, add bomb craters outside a hospital. Which is, I would say, another casualty of the fevered / ridiculous conversation around AI, because you can already do this by just taking a picture from Google Earth and dropping it into an AI image editor. I don’t buy the argument that making it easier is going to encourage more people to do it, I think anyone who wants to do it already has the tools available for very little effort. I’ve probably got another blog post in me about this.
Audio & Avatars
Google DeepMind released Lyria 3.5, an updated music generation and editing model featuring improved lyrics, vocal quality, and musicality. blog.google
This is a big improvement on the previous model; everything just sounds better. Here ⬇️ is a quick test I ran, a jazzy breakbeat instrumental; I think it sounds good, especially the drum solo at the end.
It’s available in the Flow Music app, which also got an update to Gemini Flash 3.6 for better planning and lyrics and Gemini Omni Flash for video generation. You can see them in action on a video I made for an 80s hip-hop inspired retelling of The Odyssey (on my Insta).
Of note: Suno just lost a copyright case in the German court. It’s liable to go to appeal, but not a good sign for the company.
Tavus launched PAL Maker, a platform for developers to build personalized, multimodal AI assistants. tavus.io
HeyGen added Video Podcast, which can “turn any doc, link, or idea into a two-host video show with studio scenes, multi-cam cuts, and B-roll“. instagram.com/heygen_official
Creative Tools
invideo announced Agent Two, an upgrade to its agentic creative tool, with a new feature being revealed every day.
Announced so far: Agent Intelligence, a smarter agent model/scaffold with global context, memory, and rules; Expert Agents, dedicated agents for specific roles (DoP, storyboard artist, etc), with customisable Playbooks to steer their capabilities; and input source comprehension to contextually understand any reference source.
Luma added Layers in Agents, enabling users to isolate and edit individual elements of generated or uploaded images. lumalabs.ai
OpenArt added Ad Remake, which can recreate any uploaded ad with your own product, brand, and characters. threads.com/@openart_ai
TopView launched Film Studio, with a 3D shot composer, camera controls, storyboards, and more. x.com/TopviewAIhq
Assistants
Google gave Gemini Spark more Web powers by integrating directly with Chrome, letting it use logged-in accounts and saved passwords when handling tasks. blog.google
It also expanded Spark to more Pro subscribers around the World, although not yet in the UK (sigh) or EEA.
Gemini’s macOS app added system-wide dictation and reasoning, letting users write text with their voice in any app, with optional screen context, using a keyboard shortcut. blog.google
SpaceXAI released Grok Voice Think Fast 2.0, a speech-to-speech model with improved conversational reasoning, tool use reliability, and lower latency. x.ai
Grok can now build apps and publish them to the Web. x.com/grok
Perplexity updated Spaces to Projects, introducing multi-task collaboration hubs with shared file systems and persistent project memory. perplexity.ai
Social
Snapchat added Now Playing in Snap Maps, letting users connect their Spotify account and show friends what they’re listening to in real time. newsroom.snap.com
Threads users can now chat with Meta AI in DMs, optionally including a post for context. threads.com/@threads
When originally announced this was going to be public, like Grok in X, but Meta has reacted to user concerns by moving it into DMs instead.Meta updated its Ray-Ban Display glasses, adding full Threads integration and more Instagram sharing options, along with an updated Meta AI (using the Muse Spark 1.1 model) and a limited access test of neural handwriting where users can draw letters on any surface. meta.com
LinkedIn and Snapchat are cracking down on AI slop. LinkedIn is testing an option to report any post that “Seems like AI slop”, while Snapchat says that “wholly AI-generated videos will no longer be eligible for recommendation on Spotlight”.
Honestly, I’m supportive of LinkedIn’s approach; letting creators know (privately) that their content feels inauthentic is nice encouragement to do better. Whereas I think Snapchat’s approach may be heavy-handed, because there are some genuinely talented creators working wholly in AI; I wrote about some of them for the VCCP Social Clubstack newsletter recently: Anti-slop AI creators. But every platform is entitled to its own rules.
Bonus Links
How much energy do data centers and artificial intelligence use? In a nutshell: not very much on a global scale, but concentrating too many in a limited geographical area is a problem; what they are powered by (fossil or renewable) makes a big difference; and the average person’s AI usage has negligible impact compared to other activities. This is probably too nuanced for social media.
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident. Hugging Face details the recent hack of its systems by an OpenAI model. It might be a bit technical; if so, you can ask your AI assistant to translate it for you. At it’s core, one astonishing fact: the agent understood that it was taking a test and figured out that the answer key was stored on Hugging Face’s internal servers, so to get a better score on its evaluation, it breached Hugging Face’s production infrastructure and stole the test solutions.
A quick footnote on my writing process: I’ve noticed that, since I’ve upgraded my summarisation tool to use Gemini 3.6 Flash, I’ve had to do less manual editing of the output. Even so…




