Briefing
The fast feed: breaking AI news from the global tech press, deduplicated and time-ordered.
The fast feed: breaking AI news from the global tech press, deduplicated and time-ordered.
Latest in Briefing
30 storiesZero-Shot Test-Time Canonicalization using Out-of-Distribution Scoring
Pretrained vision models often misclassify inputs that are rotated, scaled, or sheared, even though these affine transformations leave the object class unchange…
Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources
The increasing integration of distributed energy resources (DERs) is crucial for power system decarbonization, yet unlocking DERs' flexibility is challenged by …
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery
Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interaction forms.…
SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards
Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although out…
Forget to Improve: On-Device LLM-Agent Continual Learning via Budget-Curated Memory
On-device language-model agents improve by accumulating experience in retrieved memory rather than by updating weights. This memory is hard-bounded and exposed:…
Variational Inference via Entropic Transport Descent
Particle-based variational inference (ParVI) methods approximate an intractable target distribution by evolving an ensemble of interacting samples. Existing app…
KidRisk: Benchmark Dataset for Children Dangerous Action Recognition
Children are naturally energetic, and during their spontaneous activities, they often encounter potentially dangerous situations, especially when lacking parent…
Supervised Post-training of Speech Foundation Models for Robust Adaptation in Speech Deepfake Detection
Large speech foundation models have shown strong potential for speech deepfake detection, but direct fine-tuning is limited by a mismatch between self-supervise…
Offline Multi-agent Continual Cooperation via Skill Partition and Reuse
Extracting skills from multi-agent offline dataset improves learning efficiency via sharing task-invariant coordination skills among tasks. In settings where ta…
TopoCast: A Topological Fidelity Framework for Evaluating Transformer-Based Time Series Forecasting
Deep learning-based models have achieved state-of-the-art performance in Time Series Forecasting (TSF), yet their evaluation remains dominated by pointwise erro…
C3-Bench: A Context-Aware Change Captioning Benchmark
While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contex…
Optimizing Abstractive Summarization With Fine-Tuned PEGASUS
Abstractive text summarization is the technique of generating a short and concise summary comprising the salient ideas of a source text without making a subset …
Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning
Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framewor…
KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction
While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established biological pathways. We …
Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need
Risk stratification for pulmonary embolism (PE) is critical for clinical decision-making. Stratification guidelines are based on patient medical records, parame…
DualEval: Joint Model-Item Calibration for Unified LLM Evaluation
Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style preferenc…
ProvenAI: Provenance-Native Traces of Evidence in Generated Answers
Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding source meaningfully …
scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay…
Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting
While 3D Gaussian Splatting (3DGS) provides an efficient and explicit representation for novel view synthesis, enforcing stylistic coherence across viewpoints r…
Diagnosing Task Insensitivity in Language Agents
Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key source of thi…
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discrim…
RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage
Medical device recalls are a critical regulatory mechanism for protecting patient safety. The growing volume of FDA recall records presents challenges in post-r…
HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models
Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-vide…
CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only…
Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance
State-of-the-art flow models generate stunning images from text or image prompts. However, they suffer from diversity collapse when generating multiple samples …

S. 1003 Signed into Law
On Friday, June 26, 2026, the President signed into law: S. 1003, the “Lulu’s Law,” which requires the Federal Communications Commission to is…

Track total merges by adoption phase in enterprise and organization reports
Building on the AI adoption phase cohorts added to the Copilot usage metrics API, organization and enterprise reports now report the total number of pull reques…
Claude Code v2.1.195

An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 …

OpenAI's Claude Mythos competitor GPT-5.6 Sol launches under government-controlled access it calls unsustainable
OpenAI's new flagship GPT-5.6 Sol beats Anthropic's Claude Mythos 5 in coding benchmarks, but the US government is forcing a restricted rollout. OpenAI isn't ha…