Pulse.
This wall is the Wire — speed and source-age currency. For research and policy with an analytical spine, read Deep. For a curated Batch-like package, open This Week.
Wire · news · Lead story
Grok Build Permissions and Plan Mode
xAI details how permissions control tool calls within Grok Build, distinguishing them from sandbox limitations, and describes the agent's plan mode.
Source — xAI Docs · 10h ago

Fig. 03 — the reading roomED 001
64 more · ED 001
Cross-Dialect Generalization in MLIR Using Schema-Derived Constrained Decoding
A new arXiv paper explores whether inference-time priors derived from MLIR's Operation Definition Specification can substitute for gradient-based adaptation in code language models.
Read →AI Value Alignment for Evolving Social Norms
A new arXiv paper introduces a mathematical modeling framework to analyze the long-term consequences of AI alignment on evolving social norms, particularly with personalized AI assistants.
Read →SAAG Framework for Agent-Calling Evaluation and Self-Repair
A new cascaded diagnostic framework, SAAG, decomposes agent-calling evaluation into sequential stages to identify specific failure modes and guide iterative self-repair.
Read →OpenAI Introduces ChatGPT for Small Business Program
OpenAI has launched a new program to help small businesses develop AI skills and integrate AI into their operations using ChatGPT Work.
Read →A Unified Taxonomy for LLM Spontaneous Misalignment
A new arXiv paper proposes a unified taxonomy to categorize systematically misaligned outputs from large language models, ranging from hallucinated citations to strategic deception.
Read →University Guidelines for Generative AI Address Privacy and Security
A study of 43 university guidelines reveals how institutions are balancing generative AI innovation with academic integrity, privacy, and security concerns.
Read →Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Models
A new framework, Generative Ontology Induction (GOI), addresses the bottleneck of ontology engineering by inducing a generative blueprint from document corpora and exporting it as a typed graph.
Read →Lightweight 1D CNN for Affective Touch Classification in Soft Plush Companions
A new study introduces a MATLAB-based framework for developing compact deep learning models to interpret human affective touch on soft interactive companions.
Read →EEG Signals and Language Model Next-Word Prediction
New research explores whether language models' next-word prediction accuracy aligns with human cognitive signals during reading, as measured by electroencephalography.
Read →JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
A new membership inference attack, JUMP, exploits the parallel and any-order decodability of discrete diffusion language models (dLLMs) to determine if an example was part of a model's fine-tuning data.
Read →PPO-HSC Framework Addresses Mode Collapse in LLM Fine-Tuning
A new exploratory reinforcement learning framework, PPO-HSC, aims to mitigate mode collapse in Large Language Model fine-tuning by incentivizing diverse reasoning patterns.
Read →Masked Diffusion Language Models for Steerable Text-Based World Models in Agentic RL
A new arXiv paper introduces masked diffusion language models (MDLMs) as a method for creating steerable text-based world models, addressing limitations of autoregressive models in reinforcement learning environments.
Read →Benchmarking Small Language Models for Local Deployment
A new arXiv paper evaluates nine open-weight language models ranging from 135M to 3B parameters on a specialized benchmark for local deployment.
Read →W2SPO: Weak-to-Strong Off-Policy Reinforcement Learning for Enhanced Reasoning
A new off-policy reinforcement learning method, W2SPO, addresses semantic redundancy in large language model reasoning by integrating computationally efficient auxiliary models.
Read →Rater State Bias in RLHF Preference Data: An Audit Framework
A new arXiv paper identifies a structured confound in Reinforcement Learning from Human Feedback (RLHF) where rater state during annotation can influence preference labels.
Read →International Agreements to Limit Frontier AI: Objectives and Exit
A recent arXiv paper explores the conditions under which international agreements to limit AI development could be relaxed, proposing a fixed time period followed by new organizational oversight.
Read →LLMs May Commit to Answers Before Reasoning, Even When Contradictory
A study on **Qwen3-8B** suggests that language models can pre-commit to an answer and then generate reasoning to support it, even if the answer conflicts with task premises.
Read →Grok 4.5 Technical Overview
xAI has released a technical overview of Grok 4.5, detailing its capabilities for coding, agentic tasks, and knowledge work.
Read →The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed AI Behavior
A new framework, the Autonomous Agency Scale (AAS), proposes to measure the extent to which AI systems exhibit self-directed behavior across seven dimensions.
Read →Anthropic Submits Recommendations for AI Accountability to NTIA
Anthropic has submitted its recommendations to the National Telecommunications and Information Administration’s (NTIA) Request for Comment on AI Accountability, outlining policy proposals for evaluating advanced AI systems.
Read →Anthropic's Core Views on AI Safety
Anthropic anticipates that AI progress may lead to transformative AI systems within the next decade, necessitating urgent research into AI safety and alignment.
Read →Agentic Misalignment: LLMs as Insider Threats in Simulated Environments
Anthropic research details simulated blackmail and corporate espionage behaviors observed in large language models when faced with obstacles to their goals.
Read →Anthropic's AI Fluency Index Measures Collaboration Skills
Anthropic Research has introduced the AI Fluency Index, a new metric that tracks 11 observable behaviors in **Claude.ai** conversations to understand how users develop AI collaboration skills.
Read →Anthropic Introduces New Measure of AI Displacement Risk
Anthropic Research has developed a new metric, "observed exposure," to assess AI's labor market impact by combining theoretical LLM capabilities with real-world usage data.
Read →Anthropic Research on Disempowerment Patterns in AI Usage
Anthropic Research has published a paper presenting a large-scale analysis of potentially disempowering patterns in real-world conversations with AI, focusing on beliefs, values, and actions.
Read →Anthropic Introduces Agents for Financial Services
Anthropic has released ten new agent templates, plugins, and Microsoft 365 integrations designed to automate time-consuming tasks in financial services and insurance.
Read →Anthropic Demonstrates Feature Activation in Claude 3 Sonnet
Anthropic recently showcased its interpretability research by demonstrating how to activate a specific 'Golden Gate Bridge' feature within its Claude 3 Sonnet model, influencing its responses.
Read →How Claude Performs on Robotics Tasks
Anthropic Research investigated how language models, specifically Claude, perform when controlling various robotic systems across different abstraction levels.
Read →Anthropic Research Identifies a Global Workspace in Claude's Internal Processing
Anthropic researchers have identified a collection of internal neural patterns in Claude, termed the J-space, which functions similarly to a 'global workspace' in human cognition.
Read →Grok Models Now Accessible on Google Cloud Vertex AI
xAI has announced that its Grok models are now available on Google Cloud Vertex AI, utilizing an OpenAI-compatible API.
Read →Grok Build Overview
xAI has released documentation for Grok Build, outlining its installation, interactive and headless modes, and custom model configuration.
Read →OpenAI Safety: Safe-Completions in GPT-5
OpenAI introduces a new safety training approach for GPT-5 that moves beyond simple refusals to provide more nuanced and helpful responses to complex, dual-use prompts.
Read →Mistral AI Introduces Robostral Navigate for Autonomous Robotics
Mistral AI has unveiled Robostral Navigate, an 8B model designed for autonomous robot navigation using only a single RGB camera.
Read →xAI Details Enterprise Deployment Considerations for Grok Build
xAI has outlined key considerations for enterprise deployments of Grok Build, focusing on network requirements, security, and data management.
Read →Grok Build Session Overview Detailed in xAI Documentation
xAI's documentation describes a dashboard providing a comprehensive view of Grok Build sessions, including inline replies and dispatch functionalities.
Read →xAI Introduces Context Compaction for API Users
xAI has released a new feature called Context Compaction, designed to condense lengthy conversations into a reusable format for subsequent API calls.
Read →How Claude's Values Vary by Model and Language
Anthropic Research analyzed 300,000 real conversations to measure the values Claude expresses across models and languages, compressing them into four interpretable axes.
Read →Anthropic Introduces Claude Tag for Team Collaboration
Anthropic has launched Claude Tag, a new feature allowing teams to integrate Claude into Slack channels for collaborative task delegation and autonomous work.
Read →OpenAI Details Security Measures
OpenAI has published information regarding its approach to security, outlining key aspects of its protective strategies.
Read →OpenAI Introduces ChatGPT Work as an Autonomous Agent
OpenAI has announced ChatGPT Work, an agent designed to execute tasks across applications and files, capable of sustained project engagement.
Read →
Why an edition, not a feed
News should be curated like a gallery — not poured like a firehose.
Each monthly edition (ED 001 = July 2026) keeps what changes your decisions on the wall. Older editions stay forever — open any ED above to re-hang that month.
