Create
Release-by-release coverage of the AI creation stack: video, image, music, voice, design and code generation tools.
Release-by-release coverage of the AI creation stack: video, image, music, voice, design and code generation tools.
Latest in Create
30 storiesIV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt followin…
Generative Relightable Avatars
We present Generative Relightable Avatars (GRA), a person-specific method for photorealistic free-view rendering and environment-map relighting of full-body hum…
OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis
Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generated videos, they do not directly yield relia…
Cage-based Texture Transfer with Geometric Filtering
Real-time texture transfer expands the creative horizon for interactive applications, enabling seamless detail projection in scenarios that range from digital c…
Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chin…
Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization
Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available…
Semantic Browsing: Controllable Diversity for Image Generation
Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at the cost of diversity: generated samples tend…
Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation
Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acceleration methods have…
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or dependent…
Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discr…
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Comp…
VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where…
OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning
Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended t…
MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation
Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelity with r…

Build a healthcare appointment agent with Amazon Nova 2 Sonic
In this post, you will learn how to build a voice agent that handles appointment reminder conversations using Amazon Nova 2 Sonic and Amazon Bedrock AgentCore. …

Figma bets on human judgment at Config 2026 while the AI powering its canvas belongs to someone else
At Config 2026, Figma turned its canvas into a full workspace with code, animation, shaders, and AI agents. But the intelligence powering all of it is rented fr…

How Loka Built a Natural, Low-Latency Voice Agent with Amazon Nova 2 Sonic
In this post, we demonstrate the architecture and approach Loka used to solve a common frustration: robotic, slow voice assistants that cause customers to hang …
Information-Theoretic Classifier-Free Guidance with Adaptive Schedule Optimization
Diffusion models have achieved strong performance in image, text-to-image, and video generation, where conditional generation is often controlled by classifier-…
EPEdit: Redefining Image Editing with Generative AI and User-Centric Design
The demand for image manipulation has seen a significant increase recently. Traditional tools like Photoshop and Capture One, while powerful, require considerab…
DramaDirector: Geometry-Guided Short Drama Generation
Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-on…
ZONOS2 Technical Report
We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across …
S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing
We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose image gene…
CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation
Chinese news text contains dense written forms such as scores, hyphenated model names, ranges, unit symbols, percentages, English abbreviations, and mixed Chine…
DiffusionBench: On Holistic Evaluation of Diffusion Transformers
Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While methods imp…

Build a protein research copilot with Amazon Bedrock AgentCore
This post shows you how to build a conversational protein research assistant that combines three capabilities: Natural language query parsing to extract structu…

Cursor announces its own AI model, a new Git platform, and a mobile app
Cursor has revealed new details about its first AI model trained entirely in-house and announced two new products. The article Cursor announces its own AI model…

ByteDance's Seedance 2.5 breaks the 30-second barrier for AI video generation
ByteDance introduced five new AI models at Volcano Engine's FORCE conference. The centerpiece is Seedance 2.5, a video model set to launch in early July. The ar…
AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions
Agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span litera…
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain vi…
RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation
Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when…