Practice
How AI actually ships: enterprise case studies, engineering lessons from production, adoption data and ROI evidence.
How AI actually ships: enterprise case studies, engineering lessons from production, adoption data and ROI evidence.
Latest in Practice
30 storiesEnterprise-managed settings now support strictKnownMarketplaces in VS Code and GitHub Copilot CLI
Enterprises can now control which plugins their users can install in GitHub Copilot CLI and VS Code. This setting is now available in public preview. Add strict…
Copilot code review: Analysis depth and efficiency updates
Copilot code review now uses the built-in file exploration tools available in the Copilot CLI and SDK, significantly improving review cost efficiency with no ch…
UC-Search: Risk-Aware Test-Time Search for Delayed Constrained Time-Series Control
Time-series models are usually scored as forecasters, yet deployed systems often require delayed decisions under uncertainty and hard feasibility constraints. U…
PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling
Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and content cr…
Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joint deployment conditions remains insufficie…
BitNet Text Embeddings
LLM-based text embedders have substantially improved retrieval and semantic representation quality, but their deployment remains costly: large backbone models s…
AI Snitches Get Glitches: Towards Evading Agentic Surveillance
To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs. Many employers (and e…
Enhancing Brain MRI Anomaly Detection and Reasoning with ROI Rethink and Synthetic Data
Medical vision-language models typically generate diagnoses through single-pass inference without indicating which image regions support their conclusions. This…
Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem
As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent assess whether an unknow…

OpenAI and Broadcom unveil "Jalapeño," a custom chip built for LLM inference
OpenAI is adding custom hardware to its tech stack. The "Jalapeño" chip, developed with Broadcom, is tailored for large language model inference and is set to r…

OpenAI's deployment chief on Codex growth, falling AI prices, and the ROI question
OpenAI deployment chief Arnaud Fournier explains in an interview how DeployCo wants to embed AI deep inside large corporations using its own engineers. He talks…
An Introduction to Causal Reinforcement Learning
Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfac…
Entity Resolution via Batched Oracle Queries
We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity. We study how to interroga…
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference
Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting susta…
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded rea…
To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias
As Large Language Models are increasingly deployed in critical applications, robustly evaluating their social biases is paramount. However, the current literatu…
Grad Detect: Gradient-Based Hallucination Detection in LLMs
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to generating hallucinations. Detecting these…
NVIDIA and AWS Collaborate to Bring AI to Production at Scale
Building AI systems at scale is demanding, requiring low-latency inference, fast vector search, strong GPU price-performance and infrastructure that can grow wi…

OMB Advances Revolutionary FAR Overhaul with Formal Publication of Regulatory Changes Proposed changes will make common sense the hallmark of federal contracting and help the world’s largest buyer accelerate mission delivery, save money, and strengthen competition
WASHINGTON, D.C. — The Office of Management and Budget (OMB) announced the first of three releases of proposed rules to implement the Revolutionary Federal Acqu…

Introducing Mistral OCR 4
Mistral OCR 4 delivers enterprise document AI with 170-language support, bounding boxes, and self-hosted deployment.…

How Businesses Are Building Specialized AI They Can Trust
Companies are asking how to build specialized AI that fits with the way their workflows actually run. The first wave of enterprise AI was about access. Companie…
GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what th…
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking
As retrieval systems scale, high-quality reranking becomes increasingly important. However, most existing rerankers, whether encoder-based or decoder-based, joi…
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning
Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety s…
Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks
Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these developme…
Understanding Knowledge Distillation in Post-Training: When It Helps and When It Fails
Large language models (LLMs) achieve strong performance across many tasks, but their high computational cost limits deployment in resource-constrained environme…
The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production
Latent diffusion approaches to sign language production (SLP) rely on an initial stage that learns an encoding of sign pose sequences, enabling generative model…
Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning
Multi-turn tool-using agents must coordinate long-horizon tool sequences while tracking dialogue state and policy constraints. Existing approaches often separat…
The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection
Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models such as Chatterbox, pr…
AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span app…