SIGNAL
Tracking the global AI frontier — labs · research · agents · policy
Frontier Signal

Briefing

The fast feed: breaking AI news from the global tech press, deduplicated and time-ordered.

The fast feed: breaking AI news from the global tech press, deduplicated and time-ordered.

Latest in Briefing

30 stories
Optimizing Abstractive Summarization With Fine-Tuned PEGASUS
Research

Optimizing Abstractive Summarization With Fine-Tuned PEGASUS

Abstractive text summarization is the technique of generating a short and concise summary comprising the salient ideas of a source text without making a subset …

Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning
Research

Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning

Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framewor…

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction
Research

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction

While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established biological pathways. We …

Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need
Research

Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need

Risk stratification for pulmonary embolism (PE) is critical for clinical decision-making. Stratification guidelines are based on patient medical records, parame…

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation
Research

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style preferenc…

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers
Research

ProvenAI: Provenance-Native Traces of Evidence in Generated Answers

Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding source meaningfully …

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
Research

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay…

Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting
Research

Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting

While 3D Gaussian Splatting (3DGS) provides an efficient and explicit representation for novel view synthesis, enforcing stylistic coherence across viewpoints r…

Diagnosing Task Insensitivity in Language Agents
Research

Diagnosing Task Insensitivity in Language Agents

Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key source of thi…

Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
Research

Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding

Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discrim…

RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage
Research

RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage

Medical device recalls are a critical regulatory mechanism for protecting patient safety. The growing volume of FDA recall records presents challenges in post-r…

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models
Research

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-vide…

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
Research

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only…

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance
Research

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance

State-of-the-art flow models generate stunning images from text or image prompts. However, they suffer from diversity collapse when generating multiple samples …

S. 1003 Signed into Law
Policy

S. 1003 Signed into Law

On Friday, June 26, 2026, the President signed into law: S. 1003, the “Lulu’s Law,” which requires the Federal Communications Commission to is…

Track total merges by adoption phase in enterprise and organization reports
Create

Track total merges by adoption phase in enterprise and organization reports

Building on the AI adoption phase cohorts added to the Copilot usage metrics API, organization and enterprise reports now report the total number of pull reques…

Claude Code v2.1.195
Agents

Claude Code v2.1.195

An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Briefing

An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run

Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 …

OpenAI's Claude Mythos competitor GPT-5.6 Sol launches under government-controlled access it calls unsustainable
Briefing

OpenAI's Claude Mythos competitor GPT-5.6 Sol launches under government-controlled access it calls unsustainable

OpenAI's new flagship GPT-5.6 Sol beats Anthropic's Claude Mythos 5 in coding benchmarks, but the US government is forcing a restricted rollout. OpenAI isn't ha…

President Trump’s America First Agenda Scores Major Supreme Court Win on TPS Termination
Policy

President Trump’s America First Agenda Scores Major Supreme Court Win on TPS Termination

The Supreme Court has delivered a major victory for American sovereignty, ruling that the Trump Administration has full authority to terminate Temporary Protect…

Securing agentic AI with perimeter guardrails: What's new in VPC Service Controls
Practice

Securing agentic AI with perimeter guardrails: What's new in VPC Service Controls

As enterprises scale autonomous AI agents into production, enabling safe innovation requires robust architectural guardrails. AI agents connect across tools and…

MAI-Code-1-Flash for Copilot Business and Copilot Enterprise
Create

MAI-Code-1-Flash for Copilot Business and Copilot Enterprise

MAI-Code-1-Flash, Microsoft AI’s in-house coding model, is now generally available for GitHub Copilot Business and Copilot Enterprise, building on its rec…

Previewing GPT-5.6 Sol: a next-generation model
Labs

Previewing GPT-5.6 Sol: a next-generation model

OpenAI previews GPT-5.6 Sol, a next-generation model with stronger capabilities in coding, science, and cybersecurity, paired with its most advanced safety stac…

AI startup Lindy ditched Claude entirely for Deepseek, saving millions as cost pressure mounts on Anthropic
Briefing

AI startup Lindy ditched Claude entirely for Deepseek, saving millions as cost pressure mounts on Anthropic

AI startup Lindy ditched Claude entirely for Deepseek after AI costs exceeded personnel costs. CEO Flo Crivello calls it "a matter of survival for the business.…

How Cara pioneers domain-specific AI for enterprise insurance brokerages with AWS
Practice

How Cara pioneers domain-specific AI for enterprise insurance brokerages with AWS

In this post, we explore how Cara, built in cooperation with AWS, addresses these challenges. We walk through the technical design decisions and the AWS service…

Build interactive PDF text extraction from Amazon S3
Practice

Build interactive PDF text extraction from Amazon S3

In this post, you’ll build a server that extracts text from PDF files in Amazon S3 in real time. This protocol-based approach provides programmatic document acc…

Production-grade AI agents for financial compliance: Lessons from Stripe
Practice

Production-grade AI agents for financial compliance: Lessons from Stripe

In this post, you learn how Stripe built a production-grade AI agent system for financial compliance. We cover the technical architecture of Stripe’s ReAct agen…

Altman won't go public for less than $1 trillion, so OpenAI's IPO may slip to 2027
Briefing

Altman won't go public for less than $1 trillion, so OpenAI's IPO may slip to 2027

Advisors are telling OpenAI to hold off on going public until next year. The triggers: volatile tech markets and SpaceX's weak stock performance after its recor…

Anthropic doesn't need junior engineers anymore thanks to AI and warns of an economic shock when other industries follow
Briefing

Anthropic doesn't need junior engineers anymore thanks to AI and warns of an economic shock when other industries follow

"Returns on intuition": Why Anthropic no longer needs junior engineers and warns of an economic shock. The article Anthropic doesn't need junior engineers …

Building profit resilience in European private banking
Practice

Building profit resilience in European private banking

Private banks face a critical decision point: Strengthen their business focus and monetization discipline, or risk declining profits.…