SIGNAL
Tracking the global AI frontier — labs · research · agents · policy
Frontier Signal

Practice

How AI actually ships: enterprise case studies, engineering lessons from production, adoption data and ROI evidence.

How AI actually ships: enterprise case studies, engineering lessons from production, adoption data and ROI evidence.

Latest in Practice

30 stories
Enterprise-managed settings now support strictKnownMarketplaces in VS Code and GitHub Copilot CLI
Create

Enterprise-managed settings now support strictKnownMarketplaces in VS Code and GitHub Copilot CLI

Enterprises can now control which plugins their users can install in GitHub Copilot CLI and VS Code. This setting is now available in public preview. Add strict…

Copilot code review: Analysis depth and efficiency updates
Create

Copilot code review: Analysis depth and efficiency updates

Copilot code review now uses the built-in file exploration tools available in the Copilot CLI and SDK, significantly improving review cost efficiency with no ch…

UC-Search: Risk-Aware Test-Time Search for Delayed Constrained Time-Series Control
Research

UC-Search: Risk-Aware Test-Time Search for Delayed Constrained Time-Series Control

Time-series models are usually scored as forecasters, yet deployed systems often require delayed decisions under uncertainty and hard feasibility constraints. U…

PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling
Research

PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling

Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and content cr…

Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
Research

Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints

Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joint deployment conditions remains insufficie…

BitNet Text Embeddings
Research

BitNet Text Embeddings

LLM-based text embedders have substantially improved retrieval and semantic representation quality, but their deployment remains costly: large backbone models s…

AI Snitches Get Glitches: Towards Evading Agentic Surveillance
Research

AI Snitches Get Glitches: Towards Evading Agentic Surveillance

To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs. Many employers (and e…

Enhancing Brain MRI Anomaly Detection and Reasoning with ROI Rethink and Synthetic Data
Research

Enhancing Brain MRI Anomaly Detection and Reasoning with ROI Rethink and Synthetic Data

Medical vision-language models typically generate diagnoses through single-pass inference without indicating which image regions support their conclusions. This…

Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem
Research

Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem

As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent assess whether an unknow…

OpenAI and Broadcom unveil "Jalapeño," a custom chip built for LLM inference
Briefing

OpenAI and Broadcom unveil "Jalapeño," a custom chip built for LLM inference

OpenAI is adding custom hardware to its tech stack. The "Jalapeño" chip, developed with Broadcom, is tailored for large language model inference and is set to r…

OpenAI's deployment chief on Codex growth, falling AI prices, and the ROI question
Briefing

OpenAI's deployment chief on Codex growth, falling AI prices, and the ROI question

OpenAI deployment chief Arnaud Fournier explains in an interview how DeployCo wants to embed AI deep inside large corporations using its own engineers. He talks…

An Introduction to Causal Reinforcement Learning
Research

An Introduction to Causal Reinforcement Learning

Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfac…

Entity Resolution via Batched Oracle Queries
Research

Entity Resolution via Batched Oracle Queries

We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity. We study how to interroga…

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference
Research

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting susta…

AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
Research

AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning

Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded rea…

To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias
Research

To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias

As Large Language Models are increasingly deployed in critical applications, robustly evaluating their social biases is paramount. However, the current literatu…

Grad Detect: Gradient-Based Hallucination Detection in LLMs
Research

Grad Detect: Gradient-Based Hallucination Detection in LLMs

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to generating hallucinations. Detecting these…

NVIDIA and AWS Collaborate to Bring AI to Production at Scale
Labs

NVIDIA and AWS Collaborate to Bring AI to Production at Scale

Building AI systems at scale is demanding, requiring low-latency inference, fast vector search, strong GPU price-performance and infrastructure that can grow wi…

OMB Advances Revolutionary FAR Overhaul with Formal Publication of Regulatory Changes Proposed changes will make common sense the hallmark of federal contracting and help the world’s largest buyer accelerate mission delivery, save money, and strengthen competition
Policy

OMB Advances Revolutionary FAR Overhaul with Formal Publication of Regulatory Changes Proposed changes will make common sense the hallmark of federal contracting and help the world’s largest buyer accelerate mission delivery, save money, and strengthen competition

WASHINGTON, D.C. — The Office of Management and Budget (OMB) announced the first of three releases of proposed rules to implement the Revolutionary Federal Acqu…

Introducing Mistral OCR 4
Labs

Introducing Mistral OCR 4

Mistral OCR 4 delivers enterprise document AI with 170-language support, bounding boxes, and self-hosted deployment.…

How Businesses Are Building Specialized AI They Can Trust
Labs

How Businesses Are Building Specialized AI They Can Trust

Companies are asking how to build specialized AI that fits with the way their workflows actually run. The first wave of enterprise AI was about access. Companie…

GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Research

GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation

Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what th…

KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking
Research

KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking

As retrieval systems scale, high-quality reranking becomes increasingly important. However, most existing rerankers, whether encoder-based or decoder-based, joi…

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning
Research

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety s…

Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks
Research

Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks

Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these developme…

Understanding Knowledge Distillation in Post-Training: When It Helps and When It Fails
Research

Understanding Knowledge Distillation in Post-Training: When It Helps and When It Fails

Large language models (LLMs) achieve strong performance across many tasks, but their high computational cost limits deployment in resource-constrained environme…

The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production
Research

The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production

Latent diffusion approaches to sign language production (SLP) rely on an initial stage that learns an encoding of sign pose sequences, enabling generative model…

Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning
Research

Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning

Multi-turn tool-using agents must coordinate long-horizon tool sequences while tracking dialogue state and policy constraints. Existing approaches often separat…

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection
Research

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection

Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models such as Chatterbox, pr…

AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
Research

AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction

AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span app…