Deep briefs.
System cards, research, and governance with an analytical spine. The living Wire stays on /pulse; This Week packages the Batch-like read.
deep · Jul 22, 2026
Cross-Dialect Generalization in MLIR Using Schema-Derived Constrained DecodingA new arXiv paper explores whether inference-time priors derived from MLIR's Operation Definition Specification can substitute for gradient-based adaptation in code language models.
deep · Jul 22, 2026
AI Value Alignment for Evolving Social NormsA new arXiv paper introduces a mathematical modeling framework to analyze the long-term consequences of AI alignment on evolving social norms, particularly with personalized AI assistants.
deep · Jul 22, 2026
SAAG Framework for Agent-Calling Evaluation and Self-RepairA new cascaded diagnostic framework, SAAG, decomposes agent-calling evaluation into sequential stages to identify specific failure modes and guide iterative self-repair.
policy · Jul 21, 2026
OpenAI Introduces ChatGPT for Small Business ProgramOpenAI has launched a new program to help small businesses develop AI skills and integrate AI into their operations using ChatGPT Work.
deep · Jul 21, 2026
A Unified Taxonomy for LLM Spontaneous MisalignmentA new arXiv paper proposes a unified taxonomy to categorize systematically misaligned outputs from large language models, ranging from hallucinated citations to strategic deception.
policy · Jul 21, 2026
University Guidelines for Generative AI Address Privacy and SecurityA study of 43 university guidelines reveals how institutions are balancing generative AI innovation with academic integrity, privacy, and security concerns.
deep · Jul 21, 2026
Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language ModelsA new framework, Generative Ontology Induction (GOI), addresses the bottleneck of ontology engineering by inducing a generative blueprint from document corpora and exporting it as a typed graph.
deep · Jul 21, 2026
Lightweight 1D CNN for Affective Touch Classification in Soft Plush CompanionsA new study introduces a MATLAB-based framework for developing compact deep learning models to interpret human affective touch on soft interactive companions.
deep · Jul 21, 2026
EEG Signals and Language Model Next-Word PredictionNew research explores whether language models' next-word prediction accuracy aligns with human cognitive signals during reading, as measured by electroencephalography.
deep · Jul 21, 2026
JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language ModelsA new membership inference attack, JUMP, exploits the parallel and any-order decodability of discrete diffusion language models (dLLMs) to determine if an example was part of a model's fine-tuning data.
policy · Jul 21, 2026
PPO-HSC Framework Addresses Mode Collapse in LLM Fine-TuningA new exploratory reinforcement learning framework, PPO-HSC, aims to mitigate mode collapse in Large Language Model fine-tuning by incentivizing diverse reasoning patterns.
deep · Jul 21, 2026
Masked Diffusion Language Models for Steerable Text-Based World Models in Agentic RLA new arXiv paper introduces masked diffusion language models (MDLMs) as a method for creating steerable text-based world models, addressing limitations of autoregressive models in reinforcement learning environments.
policy · Jul 21, 2026
Benchmarking Small Language Models for Local DeploymentA new arXiv paper evaluates nine open-weight language models ranging from 135M to 3B parameters on a specialized benchmark for local deployment.
policy · Jul 21, 2026
W2SPO: Weak-to-Strong Off-Policy Reinforcement Learning for Enhanced ReasoningA new off-policy reinforcement learning method, W2SPO, addresses semantic redundancy in large language model reasoning by integrating computationally efficient auxiliary models.
policy · Jul 21, 2026
Rater State Bias in RLHF Preference Data: An Audit FrameworkA new arXiv paper identifies a structured confound in Reinforcement Learning from Human Feedback (RLHF) where rater state during annotation can influence preference labels.
policy · Jul 21, 2026
International Agreements to Limit Frontier AI: Objectives and ExitA recent arXiv paper explores the conditions under which international agreements to limit AI development could be relaxed, proposing a fixed time period followed by new organizational oversight.
deep · Jul 21, 2026
LLMs May Commit to Answers Before Reasoning, Even When ContradictoryA study on **Qwen3-8B** suggests that language models can pre-commit to an answer and then generate reasoning to support it, even if the answer conflicts with task premises.
deep · Jul 21, 2026
The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed AI BehaviorA new framework, the Autonomous Agency Scale (AAS), proposes to measure the extent to which AI systems exhibit self-directed behavior across seven dimensions.
research · Jul 18, 2026
Agentic Misalignment: LLMs as Insider Threats in Simulated EnvironmentsAnthropic research details simulated blackmail and corporate espionage behaviors observed in large language models when faced with obstacles to their goals.
research · Jul 21, 2026
Anthropic's AI Fluency Index Measures Collaboration SkillsAnthropic Research has introduced the AI Fluency Index, a new metric that tracks 11 observable behaviors in **Claude.ai** conversations to understand how users develop AI collaboration skills.
research · Jul 21, 2026
Anthropic Introduces New Measure of AI Displacement RiskAnthropic Research has developed a new metric, "observed exposure," to assess AI's labor market impact by combining theoretical LLM capabilities with real-world usage data.
research · Jul 21, 2026
Anthropic Research on Disempowerment Patterns in AI UsageAnthropic Research has published a paper presenting a large-scale analysis of potentially disempowering patterns in real-world conversations with AI, focusing on beliefs, values, and actions.
research · Jul 18, 2026
How Claude Performs on Robotics TasksAnthropic Research investigated how language models, specifically Claude, perform when controlling various robotic systems across different abstraction levels.
research · Jul 19, 2026
Anthropic Research Identifies a Global Workspace in Claude's Internal ProcessingAnthropic researchers have identified a collection of internal neural patterns in Claude, termed the J-space, which functions similarly to a 'global workspace' in human cognition.
policy · Jul 21, 2026
Mistral AI Introduces Robostral Navigate for Autonomous RoboticsMistral AI has unveiled Robostral Navigate, an 8B model designed for autonomous robot navigation using only a single RGB camera.
research · Jul 18, 2026
How Claude's Values Vary by Model and LanguageAnthropic Research analyzed 300,000 real conversations to measure the values Claude expresses across models and languages, compressing them into four interpretable axes.
research · Jul 18, 2026
PalmClaw: A Native On-Device Agent Framework for Mobile PhonesA new open-source agent framework, PalmClaw, enables large language model agents to run natively on mobile phones, managing sessions, memory, skills, tools, and the agent loop directly on the device.
research · Jul 18, 2026
Evaluating AI Pentesting Agents for Real-World ScenariosA new evaluation protocol shifts assessment from task completion to validated vulnerability discovery for AI pentesting agents in complex targets.
research · Jul 18, 2026
Hugging Face Survey on Self-Improvements in Modern Agentic SystemsHugging Face Daily Papers has published a survey examining the transition of self-improving autonomous agents from research prototypes to deployed systems, focusing on controllable evolution and adaptation.
research · Jul 19, 2026
GRASP: Granularity-Aware Search Policy for Agentic RAGHugging Face Daily Papers introduces GRASP, a reinforcement learning framework designed to enhance agentic retrieval-augmented generation by enabling adaptive coordination of retrieval tools.
research · Jul 18, 2026
OpenAI Introduces GPT-Red for Automated AI Red TeamingOpenAI has unveiled GPT-Red, an automated red teaming system designed to enhance AI safety and robustness through self-play.