SIGNAL
Tracking the global AI frontier — labs · research · agents · policy
Frontier Signal

Create

Release-by-release coverage of the AI creation stack: video, image, music, voice, design and code generation tools.

Release-by-release coverage of the AI creation stack: video, image, music, voice, design and code generation tools.

Latest in Create

30 stories
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
Research

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt followin…

Generative Relightable Avatars
Research

Generative Relightable Avatars

We present Generative Relightable Avatars (GRA), a person-specific method for photorealistic free-view rendering and environment-map relighting of full-body hum…

OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis
Research

OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis

Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generated videos, they do not directly yield relia…

Cage-based Texture Transfer with Geometric Filtering
Research

Cage-based Texture Transfer with Geometric Filtering

Real-time texture transfer expands the creative horizon for interactive applications, enabling seamless detail projection in scenarios that range from digital c…

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
Research

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis

While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chin…

Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization
Research

Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization

Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available…

Semantic Browsing: Controllable Diversity for Image Generation
Research

Semantic Browsing: Controllable Diversity for Image Generation

Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at the cost of diversity: generated samples tend…

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation
Research

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation

Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acceleration methods have…

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
Research

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or dependent…

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
Research

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks

Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discr…

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
Research

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Comp…

VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks
Research

VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where…

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning
Research

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning

Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended t…

MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation
Research

MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation

Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelity with r…

Build a healthcare appointment agent with Amazon Nova 2 Sonic
Practice

Build a healthcare appointment agent with Amazon Nova 2 Sonic

In this post, you will learn how to build a voice agent that handles appointment reminder conversations using Amazon Nova 2 Sonic and Amazon Bedrock AgentCore. …

Figma bets on human judgment at Config 2026 while the AI powering its canvas belongs to someone else
Briefing

Figma bets on human judgment at Config 2026 while the AI powering its canvas belongs to someone else

At Config 2026, Figma turned its canvas into a full workspace with code, animation, shaders, and AI agents. But the intelligence powering all of it is rented fr…

How Loka Built a Natural, Low-Latency Voice Agent with Amazon Nova 2 Sonic
Practice

How Loka Built a Natural, Low-Latency Voice Agent with Amazon Nova 2 Sonic

In this post, we demonstrate the architecture and approach Loka used to solve a common frustration: robotic, slow voice assistants that cause customers to hang …

Information-Theoretic Classifier-Free Guidance with Adaptive Schedule Optimization
Research

Information-Theoretic Classifier-Free Guidance with Adaptive Schedule Optimization

Diffusion models have achieved strong performance in image, text-to-image, and video generation, where conditional generation is often controlled by classifier-…

EPEdit: Redefining Image Editing with Generative AI and User-Centric Design
Research

EPEdit: Redefining Image Editing with Generative AI and User-Centric Design

The demand for image manipulation has seen a significant increase recently. Traditional tools like Photoshop and Capture One, while powerful, require considerab…

DramaDirector: Geometry-Guided Short Drama Generation
Research

DramaDirector: Geometry-Guided Short Drama Generation

Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-on…

ZONOS2 Technical Report
Research

ZONOS2 Technical Report

We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across …

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing
Research

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing

We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose image gene…

CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation
Research

CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation

Chinese news text contains dense written forms such as scores, hyphenated model names, ranges, unit symbols, percentages, English abbreviations, and mixed Chine…

DiffusionBench: On Holistic Evaluation of Diffusion Transformers
Research

DiffusionBench: On Holistic Evaluation of Diffusion Transformers

Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While methods imp…

Build a protein research copilot with Amazon Bedrock AgentCore
Practice

Build a protein research copilot with Amazon Bedrock AgentCore

This post shows you how to build a conversational protein research assistant that combines three capabilities: Natural language query parsing to extract structu…

Cursor announces its own AI model, a new Git platform, and a mobile app
Briefing

Cursor announces its own AI model, a new Git platform, and a mobile app

Cursor has revealed new details about its first AI model trained entirely in-house and announced two new products. The article Cursor announces its own AI model…

ByteDance's Seedance 2.5 breaks the 30-second barrier for AI video generation
Briefing

ByteDance's Seedance 2.5 breaks the 30-second barrier for AI video generation

ByteDance introduced five new AI models at Volcano Engine's FORCE conference. The centerpiece is Seedance 2.5, a video model set to launch in early July. The ar…

AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions
Research

AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions

Agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span litera…

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
Research

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain vi…

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation
Research

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when…