Create
Release-by-release coverage of the AI creation stack: video, image, music, voice, design and code generation tools.
Release-by-release coverage of the AI creation stack: video, image, music, voice, design and code generation tools.
Latest in Create
30 storiesBuild Production-Ready Agents with the GitHub Copilot Harness and Agent Framework
Developers increasingly want to build agents that can reason about code, modify files, execute commands, interact with developer tools, and work across entire r…
How we built a realtime system for responsive voice AI in six months
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.…

Building agent teams with Agent Framework, GitHub Copilot CLI and Squad
Microsoft Agent Framework supports creating agents that use the GitHub Copilot SDK as their backend. GitHub Copilot agents provide access to powerful coding-ori…

Opus 5 >> Fable 5
Codex plus Voice is the new openclaw…

Best practices for applying Amazon Bedrock Guardrails to code generation workflows
In this post, we explain how Amazon Bedrock Guardrails can be configured for code generation workflows with coding assistants to overcome these constraints. Wit…
python-github-copilot-1.0.0
How Codex became a collaborator for OpenAI’s creative team
How OpenAI’s creative team uses Codex to build custom creative tools, accelerate ideation, and prototype faster with context-aware AI.…
Introducing OpenAI Presence
Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for customer and internal workf…

Building a restaurant telephony AI host with Amazon Bedrock AgentCore and Amazon Nova 2 Sonic
In this post, we show you how to build a voice ordering system that answers a phone number and takes the order from greeting to confirmation. The system uses Am…
How Cars24 scales conversations and builds faster with OpenAI
Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams acr…

ScienceSoft’s HIPAA-compliant AI voice scheduler built on AWS
In this post, you will learn how ScienceSoft, an Amazon Web Services (AWS) Services Partner, integrated Amazon Nova 2 Sonic with Amazon Bedrock Guardrails to bu…

Securing the AI supply chain on GKE: Introducing k8s-aibom for automated AI BOMs
How should your security team manage shadow AI? Workloads deployed by developers without formal registration can often evade traditional security scanners, beca…
The AI Picbreeder Experiment: Can AI agents be creative when nobody tells them what to create?
How OpenAI delivers low-latency voice AI at scale
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.…
How Deutsche Telekom is rewiring telecommunications with AI
How Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, network operations, and the future of voice.…
GPT-5.6 is now the preferred model in Microsoft 365 Copilot
Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, higher-quality work.…

Grok x Cursor
More Fable, GPT-5.6 and new ChatGPT voice…
Introducing GPT-Live
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.…

Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
Meta's Superintelligence Labs ships Muse Image, its first image generation model. Like OpenAI's GPT Image 2, it works as an agent, using tools like code executi…
Responses are the Easy Part: What We’ve Learned Building Real-time Voice Experiences at Scale
Timing, interruption, silence and recovery shape the experience as much as words…

Copilot goes cheap as Microsoft phases out OpenAI and Anthropic models to cut costs
Microsoft is replacing AI models from OpenAI and Anthropic with its own MAI models in products like Excel and Outlook. Tens of thousands of queries per week alr…

Microsoft follows Anthropic and OpenAI into the AI super app race with overhauled Copilot and AutoPilot agents
Microsoft reportedly plans to merge its consumer and enterprise Copilot apps into a single app in August. Rarely used features like Copilot Podcasts are getting…

Chinese AI video maker Kling raises $2 billion as it gears up for Hong Kong IPO
Kuaishou has raised about $2 billion from investors for its AI video division, Kling. The article Chinese AI video maker Kling raises $2 billion as it gears up …

Google launches Nano Banana 2 Lite for fast AI images and Gemini Omni Flash for video via API
Google adds two new generative AI models. Nano Banana 2 Lite generates images in four seconds at $0.034 a pop. Gemini Omni Flash brings video generation and edi…

Bringing speed and strong cost performance to the market with Gemini Omni Flash and Nano Banana 2 Lite
Great creative happens when your tools move at the speed of your ideas. To help you create rich, reliable experiences while reducing regeneration time and costs…
Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting
Recent advances in 3D Gaussian Splatting have demonstrated unprecedented success in novel view synthesis. However, the substantial inference and storage overhea…
OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data
Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Des…
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semantic align…
UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception
Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches …
ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation witho…