arXiv AI
Publisher updates
Follow the original publisher. These items are discovery material—not automatically approved reporting, endorsements, or verified breaking news from TLB.
Collection recently observed.Source timestamps describe the item, not necessarily a new event. 6014 stored items match this view.
arXiv AI
Demystifying the Privacy-Utility Trade-off in LLM Interactions
External source · read at the original publisherarXiv AI
Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks
External source · read at the original publisherarXiv AI
The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures
External source · read at the original publisherarXiv AI
Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents
External source · read at the original publisherarXiv AI
Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning
External source · read at the original publisherarXiv AI
MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG
External source · read at the original publisherarXiv AI
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
External source · read at the original publisherarXiv AI
KuaiRP Series Role-playing Models Technical Report
External source · read at the original publisherarXiv AI
Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment
External source · read at the original publisherarXiv AI
The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems
External source · read at the original publisherarXiv AI
Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation
External source · read at the original publisherarXiv AI
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat
External source · read at the original publisherarXiv AI
Breaking Predictions Is Not Enough: Specified-Foil Counterfactuals for Temporal Graphs
External source · read at the original publisherarXiv AI
Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation
External source · read at the original publisherarXiv AI
SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics
External source · read at the original publisherarXiv AI
Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment
External source · read at the original publisherarXiv AI
Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce
External source · read at the original publisherarXiv AI
An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning
External source · read at the original publisherarXiv AI
CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series Forecasting
External source · read at the original publisherarXiv AI
A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies
External source · read at the original publisherarXiv AI
NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment
External source · read at the original publisherarXiv AI
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
External source · read at the original publisherarXiv AI
AI-Powered Flare Combustion Efficiency Estimation
External source · read at the original publisherarXiv AI
Predicting Train Delays in Finland Using Machine Learning and Weather Data
External source · read at the original publisherarXiv AI
Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1
External source · read at the original publisherarXiv AI
When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting
External source · read at the original publisherarXiv AI
Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data
External source · read at the original publisherarXiv AI
Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model
External source · read at the original publisherarXiv AI