Create
Release-by-release coverage of the AI creation stack: video, image, music, voice, design and code generation tools.
Release-by-release coverage of the AI creation stack: video, image, music, voice, design and code generation tools.
Latest in Create
30 storiesThe Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection
Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models such as Chatterbox, pr…
Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping
Reference-based diffusion models enable highly controllable image generation by leveraging elements from input images to guide prompt-driven synthesis. However,…

Running ComfyUI workflows on Amazon SageMaker AI processing jobs
In this post, we walk you through how to deploy ComfyUI workflows on Amazon SageMaker AI processing jobs to generate hundreds of high-quality images in a single…
LooseControlVideo: Directorial Video Control using Spatial Blocking
Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and tem…

Amazon drops its OpenAI drama film after signing a $50 billion deal with Sam Altman's company
Amazon MGM Studios has dropped "Artificial," the nearly finished OpenAI film directed by Luca Guadagnino with Andrew Garfield as Sam Altman. Amazon struck a $50…

Adobe adds AI agents to Photoshop, Premiere, and more Creative Cloud apps
Adobe is rolling out its "creative agent" across its main Creative Cloud apps and third-party AI platforms like ChatGPT and Claude. Users describe what they wan…

Midjourney, known for AI image generation, unveils a full-body ultrasound scanner and its own spa
Rumors about Midjourney hardware have circulated for years, but nobody saw this coming. The AI image startup is building a full-body ultrasound scanner and open…
UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation
Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. Howeve…

WHAT THEY ARE SAYING: First Lady Melania Trump Launches Fostering the Future Accounts
Last week, First Lady Melania Trump launched Fostering the Future Accounts, a new financial resource established to empower foster youth to be fiscally aut…

Microsoft's Copilot Cowork moves to usage-based billing and may tap DeepSeek
Microsoft is weighing a fine-tuned version of Deepseek V4 as a cheaper model option for Copilot Cowork. The company is also switching to usage-based billing, si…

SpaceX bets $60 billion on Cursor to catch OpenAI and Anthropic
Just two trading days after its IPO, Elon Musk's SpaceX is buying AI coding startup Anysphere. The deal is meant to help the struggling xAI division catch up wi…

Microsoft Research's Mirage gives video generation a persistent spatial memory that doesn't forget what's around the corner
Mirage, a video world model from Microsoft Research and several universities, stores scene information directly in latent space instead of pixel-based point clo…

Announcing Ironwood TPUs General Availability and new Axion VMs to power the age of inference
Today’s frontier models, including Google’s Gemini, Veo, Imagen, and Anthropic’s Claude train and serve on Tensor Processing Units (TPUs). For many organization…

Apply AI Webinar on AI for Cultural, Creative and Media Sectors
Apply AI Webinar on AI for Cultural, Creative and Media Sectors Anonymous (not verified) Tue, 06/02/2026 - 12:29 24 June 2026 This webinar will focus on …

Nano Banana 2 and Nano Banana Pro are generally available, and already powering creative workflows
Organizations are unlocking entirely new ways to use image generation and editing across their industries. To drive next-generation experiences, businesses are …

Developer's guide to Gemini Enterprise and A2UI integration
If you've built a chatbot, you know this conversation: User: "Book a table for two tomorrow at 7pm." Agent: "Okay, for what day?" User: "Tomorrow." Agent: "What…

Doubling Down on Suno, the Platform for Creative Entertainment
Every major consumer platform is built on a new behavior. TikTok made short-form video consumption mainstream. Netflix changed how we watch TV. Suno is doing so…

Evaluate your Amazon Nova Sonic voice agent at scale, no microphone required
In this post, we walk you through the Nova Sonic Test Harness, an open source framework that we built to solve both problems. It serves as a rapid iteration too…

Modernizing Healthcare: How Alcidion achieved greater stability and performance with AlloyDB
In clinical informatics, every second counts. For Alcidion, a global leader in smart health solutions, the mission is simple but critical: use technology to red…

It’s safe to close your laptop now: Hosting coding agents on Amazon Bedrock AgentCore
Amazon Bedrock AgentCore Runtime gives each agent session its own isolated microVM with a persistent workspace, secure tool access through Gateway, and built-in…

MCP Apps - Bringing UI Capabilities To MCP Clients
MCP Apps is now live as the first official MCP extension — tools can return interactive UI components that render directly in the conversation.…

Whole-Body Conditioned Egocentric Video Prediction
.modal { display: none; position: fixed; z-index: 9999; padding-top: 50px; left: 0; top: 0; width: 100%; height: 100%; overflow: auto; background-color: rgba(0,…

GridSFM: A new, small foundation model for the electric grid
Introducing GridSFM, a small foundation model that can predict AC optimal power flow in milliseconds, boosting efficiency and unlocking cost savings. Learn how …
TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment
Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on…
Limits of spectral learning under noise
Learning functional relationships from noisy data is a central problem in scientific inference. Spectral methods approximate unknown functions by expanding them…
DuET: Dual Expert Trajectories for Diffusion Image Editing
Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent source-image con…
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying…
Modality Forcing for Scalable Spatial Generation
Text-to-image (T2I) models contain rich spatial priors. Synthesizing photorealistic, cluttered scenes requires an understanding of geometry, including perspecti…
InterleaveThinker: Reinforcing Agentic Interleaved Generation
Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constr…
A new way to express yourself: Gemini can now create music
The Gemini app now features our most advanced music generation model Lyria 3, empowering anyone to make 30-second tracks using text or images.…