SIGNAL
Tracking the global AI frontier — labs · research · agents · policy
Frontier Signal

Practice

How AI actually ships: enterprise case studies, engineering lessons from production, adoption data and ROI evidence.

How AI actually ships: enterprise case studies, engineering lessons from production, adoption data and ROI evidence.

Latest in Practice

30 stories
From Local to Production: Deploy Your Microsoft Agent Framework Agent with Foundry Hosted Agents
Agents

From Local to Production: Deploy Your Microsoft Agent Framework Agent with Foundry Hosted Agents

Once you have your Microsoft Agent Framework (MAF) agent or workflow happily running locally on your dev machine, it’s time to decide how to deploy your a…

Governance at the Speed of Agents: Microsoft Agent Framework and Agent Governance Toolkit, Better Together
Agents

Governance at the Speed of Agents: Microsoft Agent Framework and Agent Governance Toolkit, Better Together

Building powerful AI agents is only half the story, running them safely in production is the real challenge. As customers adopt Microsoft Agent Framework for ag…

Stop prompt injection from hijacking your agent, new security capabilities now released within Agent Framework
Agents

Stop prompt injection from hijacking your agent, new security capabilities now released within Agent Framework

Prompt injection is the #1 risk on the OWASP LLM Top 10, and most agents in production today defend against it with one of two heuristics: a defensive system pr…

ICYMI: Inside the Microsoft Agent Framework: How we designed a layered SDK
Agents

ICYMI: Inside the Microsoft Agent Framework: How we designed a layered SDK

In case you missed it, the Command Line blog was launched last week and has a great article (by yours truly) about our SDK design philosophy with Microsoft Agen…

Scaling Up Reinforcement Learning for Traffic Smoothing: A 100-AV Highway Deployment
Research

Scaling Up Reinforcement Learning for Traffic Smoothing: A 100-AV Highway Deployment

Training Diffusion Models with Reinforcement Learning We deployed 100 reinforcement learning (RL)-controlled cars into rush-hour highway traffic to smooth conge…

Identifying Interactions at Scale for LLMs
Research

Identifying Interactions at Scale for LLMs

--> Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial inte…

Notes from inside China's AI labs
Research

Notes from inside China's AI labs

Lessons from my trip to talk to most of the leading AI labs in China.…

Building realistic electric transmission grid dataset at scale: a pipeline from open dataset
Research

Building realistic electric transmission grid dataset at scale: a pipeline from open dataset

Microsoft Research is excited to release an open dataset of approximate transmission topology of the U.S. power grid derived from publicly available data. The a…

MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models
Research

MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models

MagenticLite is an agentic system for small models that works across the browser and local file system in a single workflow. It combines specialized models and …

Data Formulator 0.7: AI-powered data analytics for enterprise data
Research

Data Formulator 0.7: AI-powered data analytics for enterprise data

Data Formulator introduces AI-powered analytics for enterprise data workflows. Data teams can easily bring enterprise data into an AI-ready workspace where user…

TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization
Research

TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization

Large language models (LLMs) exhibit exceptional general language processing capabilities, but their memory and compute costs hinder deployment. Ternarization h…

Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models
Research

Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models

Large language models are increasingly deployed in citation-augmented settings, yet the effect of citation presence on model behavior independent of factual con…

MiniPIC: Flexible Position-Independent Caching in <100LOC
Research

MiniPIC: Flexible Position-Independent Caching in <100LOC

Retrieval-augmented and agentic workloads repeatedly prefill recurring predictable structured inputs (which we call "spans") such as documents and code files. Y…

OR-Action: Multi-Role Video Understanding with Fine-Grained Actions
Research

OR-Action: Multi-Role Video Understanding with Fine-Grained Actions

Fine-grained understanding of operating room (OR) activity could enable workflow-aware assistance, yet remains difficult due to clutter, occlusions, and limited…

From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent
Research

From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent

Large language models (LLMs) have shown promise in automating scientific peer review. However, existing approaches often struggle to generate in-depth reviews s…

Uncertainty Estimation for Molecular Diffusion Models
Research

Uncertainty Estimation for Molecular Diffusion Models

Diffusion models have seen wide adoption for 3D molecular generation, yet they offer no principled signal of when a generated molecule is likely to be of low qu…

Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch
Research

Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch

Dispatch in three-sided marketplaces provides a natural setting for reinforcement learning from world feedback: decisions are evaluated by delayed operational o…

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
Research

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, …

OpenAI announces Frontier Alliance Partners
Labs

OpenAI announces Frontier Alliance Partners

OpenAI announces Frontier Alliance Partners to help enterprises move from AI pilots to production with secure, scalable agent deployments.…

OpenAI Codex and Figma launch seamless code-to-design experience
Labs

OpenAI Codex and Figma launch seamless code-to-design experience

OpenAI and Figma launch a new Codex integration that connects code and design, enabling teams to move between implementation and the Figma canvas to iterate and…

Nano Banana 2: Combining Pro capabilities with lightning-fast speed
Labs

Nano Banana 2: Combining Pro capabilities with lightning-fast speed

Our latest image generation model offers advanced world knowledge, production ready specs, subject consistency and more, all at Flash speed.…

OpenAI and Amazon announce strategic partnership
Labs

OpenAI and Amazon announce strategic partnership

OpenAI and Amazon announce a strategic partnership bringing OpenAI’s Frontier platform to AWS, expanding AI infrastructure, custom models, and enterprise AI age…

Our agreement with the Department of War
Labs

Our agreement with the Department of War

Details on OpenAI’s contract with the Department of War, outlining safety red lines, legal protections, and how AI systems will be deployed in classified enviro…

Gemini 3.1 Flash-Lite: Built for intelligence at scale
Labs

Gemini 3.1 Flash-Lite: Built for intelligence at scale

Gemini 3.1 Flash-Lite is our fastest and most cost-efficient Gemini 3 series model yet.…

How Axios uses AI to help deliver high-impact local journalism
Labs

How Axios uses AI to help deliver high-impact local journalism

Axios COO Allison Murphy explains how the company uses AI to support local reporters, streamline newsroom workflows, and deliver high-impact local journalism at…

Introducing the Adoption news channel
Labs

Introducing the Adoption news channel

Practical insights and frameworks to turn AI progress into business advantage…

How Descript engineers multilingual video dubbing at scale
Labs

How Descript engineers multilingual video dubbing at scale

Using OpenAI reasoning models, Descript unlocked automatic localization of large content libraries without losing timing or meaning.…

Wayfair boosts catalog accuracy and support speed with OpenAI
Labs

Wayfair boosts catalog accuracy and support speed with OpenAI

Wayfair uses OpenAI models to improve ecommerce support and product catalog accuracy, automating ticket triage and enhancing millions of product attributes at s…

How we monitor internal coding agents for misalignment
Labs

How we monitor internal coding agents for misalignment

How OpenAI uses chain-of-thought monitoring to study misalignment in internal coding agents—analyzing real-world deployments to detect risks and strengthen AI s…

Accelerating the next phase of AI
Labs

Accelerating the next phase of AI

OpenAI raises $122 billion in new funding to expand frontier AI globally, invest in next-generation compute, and meet growing demand for ChatGPT, Codex, and ent…