LIVE · refreshes every 20 min
updated Sep 16, 04:23 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
☾
110010
.art
/
archive
/
research
/ 2026-07
Research — July 2026
Jul 31
Evidence-ledger adjudication for claim-evidence traceability in AI agents assesses support relations and routes unsupported or contradicted claims back to the author
arXiv cs.AI
source ↗
Jul 31
Latent channels in multi-agent LLMs are evaluated for actual communication through a causal audit of latent messages.
arXiv cs.AI
source ↗
Jul 31
EvoPINN: agentic discovery of executable algorithms for physics-informed neural networks
arXiv cs.AI
source ↗
Jul 31
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
arXiv cs.AI
source ↗
Jul 31
UrbanDS: a graph-guided LLM multi-agent system for data-intensive urban tasks
arXiv cs.AI
source ↗
Jul 31
Objective misalignment in mixed-motive LLM multi-agent systems evaluated with Werewolf game
arXiv cs.AI
source ↗
Jul 31
Rethinking self-evolution: a constrained exploration-exploitation process for mitigating skill overfitting
arXiv cs.AI
source ↗
Jul 31
AI agents discover statistical mechanical mappings from a raw partition function to a tractable representation in six Ising-type problems; StatMechBench-v0 benchmark introduced
arXiv cs.AI
source ↗
Jul 31
GuideSkill: evolving executable LLM agent skills for guideline-grounded clinical reasoning
arXiv cs.AI
source ↗
Jul 31
MultivationBench: a benchmark for multimodal sequential motivation reasoning
arXiv cs.AI
source ↗
Jul 31
ClinLens: a benchmark of 200 executable tasks over five linked MIMIC resources for longitudinal multimodal clinical data science
arXiv cs.AI
source ↗
Jul 31
CG-World: a large-scale world-state dataset and protocol for world models
arXiv cs.AI
source ↗
Jul 31
Passive video to editable experience: Pegasus translates human demonstrations into robot-learnable data through structured knowledge transfer
arXiv cs.AI
source ↗
Jul 31
Representational quality accounts for RL models’ superior mathematical reasoning performance over SFT fine-tuned models
arXiv cs.AI
source ↗
Jul 31
Evaluation scores are perishable knowledge claims
arXiv cs.AI
source ↗
Jul 31
When benchmark inferences do not compose: projectibility in AI evaluation
arXiv cs.AI
source ↗
Jul 31
TraceCoder enables explainable and auditable code generation with position-key snippet versioning
arXiv cs.AI
source ↗
Jul 31
GoGoTB: agentic RTL verification with specification-grounded coverage closure
arXiv cs.AI
source ↗
Jul 31
What Does It Take to Detect an AI Agent? Minimal feature sets for behavioral detection under browser automation
arXiv cs.AI
source ↗
Jul 31
Eco3S: a socio-economic system simulation framework for agent-based modeling and policy analysis
arXiv cs.AI
source ↗
Jul 31
AlphaSchema: exploring the space of trading semantics for LLM-based alpha mining
arXiv cs.AI
source ↗
Jul 31
Property-driven causal abstractions for Markov decision processes
arXiv cs.AI
source ↗
Jul 31
Belief-guided decision making with uncertainty gating in the game of Go
arXiv cs.AI
source ↗
Jul 31
Fewer clarifications, better code: benchmarking cross-session personalized ambiguity adaptation in coding assistants
arXiv cs.AI
source ↗
Jul 31
ICLE++: modeling fine-grained traits for holistic essay scoring
arXiv cs.CL
source ↗
Jul 31
Systematic evaluation of 41 open-weight language models for zero-shot intent classification across eight datasets (135M–9B parameters) and 15 model families
arXiv cs.CL
source ↗
Jul 31
Harness-G: a graph-structured harness for search agents
arXiv cs.CL
source ↗
Jul 31
Kinetics of training describe a driven-nucleation rate law for emergence, plasticity loss, and circuit control in language models
arXiv cs.LG
source ↗
Jul 31
From single- to cross-document: benchmarking multi-granularity event analysis of large language models
arXiv cs.CL
source ↗
Jul 31
Rethinking EEG-based disease diagnosis: decoupling instance representation learning from subject-level supervision
arXiv cs.LG
source ↗
Jul 31
Grounded agentic extraction and expert-adjudicated evaluation of intertextuality in classical Chinese histories
arXiv cs.CL
source ↗
Jul 31
Recursive transformers for semiconductor thermo-mechanical reliability
arXiv cs.LG
source ↗
Jul 31
Modeling decisions in blockchain analytics: a leakage-aware evaluation of tree-based vs. sequential models
arXiv cs.LG
source ↗
Jul 31
SE(3)-MeanFlow: few-step protein backbone generation on Lie groups
arXiv cs.LG
source ↗
Jul 31
AHA-Memes: a fine-grained multimodal benchmark for understanding hate in Arabic memes
arXiv cs.CL
source ↗
Jul 31
FunL2O: LLM-guided feature function design for learning to optimize
arXiv cs.LG
source ↗
Jul 31
THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model
arXiv cs.LG
source ↗
Jul 31
Good rankers, bad objectives: bilinear contrastive critics under expressive policy search
arXiv cs.LG
source ↗
Jul 31
DualAnchor: preserving language priors and improving lexical fidelity in gloss-free sign language translation
arXiv cs.CL
source ↗
Jul 31
Context-informed ship trajectory prediction via conditional attention
arXiv cs.LG
source ↗
Jul 31
ZUNA1.1: a 380M-parameter diffusion autoencoder for flexible EEG reconstruction across variable lengths, channels, and intervals
arXiv cs.LG
source ↗
Jul 31
DoTime: a synthetic benchmark generator for interventional and counterfactual time series
arXiv cs.LG
source ↗
Jul 31
AI-assisted pre-review of open-source submissions used at BOSC 2026 to help volunteer reviewers by pre-reviewing abstracts for certain criteria
arXiv cs.CL
source ↗
Jul 31
BridgeAlign: Bridging preference alignment for humanities and social sciences
arXiv cs.CL
source ↗
Jul 31
Sympathetic framing: evaluating AI alignment across sociodemographic groups
arXiv cs.CL
source ↗
Jul 31
Prompt chaining in practice: a case study in automated scholarly report generation
arXiv cs.CL
source ↗
Jul 31
LVLMs evaluated on perception and reasoning jointly to uncover truth behind visual illusions
arXiv cs.CL
source ↗
Jul 31
Same facts, different diagnosis: measuring and mitigating narrative anchoring in clinical language models
arXiv cs.CL
source ↗
Jul 31
Gradient-free task-conditioned retrieval for on-device in-context learning
arXiv cs.CL
source ↗
Jul 31
Flat Score, Amplified Failures: Quantization to 4-bit weights hides damage in multi-turn tool-calling LLM agents
arXiv cs.LG
source ↗
Jul 31
SkillSmith: learning to compose parametric skills and textual knowledge
arXiv cs.CL
source ↗
Jul 31
B1ade: 335M embedding model and 1B parameter small language models for minimalist RAG
arXiv cs.CL
source ↗
Jul 31
RLPF: reinforcement learning from performance feedback for code generation
arXiv cs.LG
source ↗
Jul 31
Compression-based behavioral similarity enables open-world Sybil discovery on Ethereum
arXiv cs.LG
source ↗
Jul 31
TIER-MoE: trust-informed expert routing via conditional modality risk for multimodal fusion in biomedical classification
arXiv cs.LG
source ↗
Jul 31
Benchmarking LLM competence on logical inference over probability operators
arXiv cs.CL
source ↗
Jul 31
Adam converges under heavy-tailed noise in the plain vector-form setting, with guarantees for stochastic gradients having bounded p-th moment (p in (1,2]).
arXiv cs.LG
source ↗
Jul 31
Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention
arXiv cs.CL
source ↗
Jul 31
ECG-InterpBench benchmarks the interpretability of ECG foundation-model representations using matched-scale sparse autoencoders
arXiv cs.LG
source ↗
Jul 31
Explorative modeling: unlocking a third pretraining axis and end-to-end generation
arXiv cs.LG
source ↗
Jul 31
Structure-aware data organization for efficient LLM post-training
arXiv cs.LG
source ↗
Jul 31
Training skills like parameters via self-supervised semantic diffusion
arXiv cs.CL
source ↗
Jul 31
Benchmarking the residual: long-horizon evaluations reveal degradation beyond short-task performance
arXiv cs.LG
source ↗
Jul 31
AWARE-FX: an auditable AI/NLP decision-support system converts corporate annual report text into traceable foreign-exchange hedging disclosure measures
arXiv cs.CL
source ↗
Jul 30
Daniela Rus receives the Bavarian Minister-President's High-Tech Prize.
MIT News (AI)
source ↗
Jul 30
Capitol Hill hosts Congressional Visit Days as researchers connect with policy makers; participants pose on the Capitol steps with Senator Alex Padilla.
MIT News (AI)
source ↗
Jul 30
Echoverse: deep, evolving environments for computer-use agents
Microsoft Research
source ↗
Jul 30
Examining the efficacy of graph neural network message-passing in regression contexts
arXiv cs.LG
source ↗
Jul 30
Automorphism-induced non-canonicity in top-k explanations of graph neural networks
arXiv cs.LG
source ↗
Jul 30
Misalignment has a personality: a Big Five account of emergent misalignment
arXiv cs.CL
source ↗
Jul 30
Characterizing human-likeness in AI generated poetry: a zero-shot classification study
arXiv cs.CL
source ↗
Jul 30
MeRLa: Meta-learned reward shaping for reinforcement learning from human feedback
arXiv cs.LG
source ↗
Jul 30
WikiLoop jointly learns to build and navigate an agent-native wiki with downstream feedback
arXiv cs.CL
source ↗
Jul 30
Primary source headline: Do Methods Support the Claims? Intra-Paper Verification for Peer Review Rewritten headline: Do methods support the claims? Intra-paper verification for peer review
arXiv cs.CL
source ↗
Jul 30
Corpus of transcribed English-language religious radio broadcasts from 785 webstreams captured over July 2025 in the United States
arXiv cs.CL
source ↗
Jul 30
High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption
arXiv cs.LG
source ↗
Jul 30
Emergent sparsity in frozen random CNN feature extractors for deep reinforcement learning.
arXiv cs.LG
source ↗
Jul 30
Mergeable model-side aggregation states for long-context language models
arXiv cs.CL
source ↗
Jul 30
Shared SFT lessons across alignment, model organisms, and toy models show transferable findings through cross-domain tests
arXiv cs.LG
source ↗
Jul 30
Between gradient and natural gradient: a continuum of LoRA initializations
arXiv cs.LG
source ↗
Jul 30
RAGuard introduces a layered defense against factual corpus-poisoning attacks in retrieval-augmented generation systems.
arXiv cs.LG
source ↗
Jul 30
Diagnosing fine-grained inconsistency classification in financial disclosure text
arXiv cs.CL
source ↗
Jul 30
DuplexGen enables adaptive synthesis of human-AI turn-taking dialogues.
arXiv cs.CL
source ↗
Jul 30
Voice Memory enables agentic speech recognition with a frozen corrector deciding per utterance to act on the hypothesis or abstain, and an asynchronous optimizer revising a per-domain memory.md through bounded edits
arXiv cs.CL
source ↗
Jul 30
Self-serve entity resolution pipeline shows no single winner across datasets; recommend multiple algorithm families and automatic bake-off selection
arXiv cs.LG
source ↗
Jul 30
ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling
arXiv cs.LG
source ↗
Jul 30
DHRCL: training code LLMs with dense hierarchical rewards and curriculum learning
arXiv cs.LG
source ↗
Jul 30
Learning implicit causal world models from multi-agent demonstrations
arXiv cs.LG
source ↗
Jul 30
Data fusion and contrastive alignment for unconstrained IR molecular structure elucidation
arXiv cs.LG
source ↗
Jul 30
Knowledge before reasoning: EC-Reason-Bench, a training-free diagnostic benchmark for LLM enzyme classification
arXiv cs.CL
source ↗
Jul 30
MetaKoopman: Bayesian meta-learning of Koopman operators for modeling nonlinear dynamics under distribution shifts
arXiv cs.LG
source ↗
Jul 30
Top-k Pareto bandits: hypervolume regret for multi-objective slate selection
arXiv cs.LG
source ↗
Jul 30
Large-scale chatbot validation using synthetic customer agents as digital twins generated from real transactional and conversational data
arXiv cs.CL
source ↗
Jul 30
SCOUT: Per-context reset curricula for sparse-reward reinforcement learning
arXiv cs.LG
source ↗
Jul 30
AgentGUI: an interface for observing and steering long-running AI agents
arXiv cs.CL
source ↗
Jul 30
Dynamic parameterization is not dynamic inference
arXiv cs.LG
source ↗
Jul 30
Choosing where and how to moderate: end-to-end trade-offs in filter placement and response rewriting
arXiv cs.CL
source ↗
Jul 30
Existence-field diffusion model for spatial point processes with variable cardinality
arXiv cs.LG
source ↗
Jul 30
Where detectors fail: expert-guided mutual distillation improves robustness to tail-domain gaps
arXiv cs.CL
source ↗
Jul 30
Sim2Win: a team-agnostic, event-based pre-match outcome prediction and tactical profiling system for football
arXiv cs.LG
source ↗
Jul 30
Early verdicts, better budgets: sequential adaptive rollout allocation for compute-efficient RLVR
arXiv cs.LG
source ↗
Jul 30
Post-training at the edge of detectability: a game-theoretic approach to fine-tuning
arXiv cs.LG
source ↗
Jul 30
Symphony-Bias investigates gender associations with musical instruments in multimodal LLMs
arXiv cs.CL
source ↗
Jul 30
FloDR: An invertible dimensionality reduction method based on a normalising flow
arXiv cs.LG
source ↗
Jul 30
When synthetic users fail: a cross-domain benchmark of LLM-simulated human survey responses
arXiv cs.CL
source ↗
Jul 30
Lossy verification in speculative decoding can rewrite the decoding distribution and affect stability, study finds
arXiv cs.CL
source ↗
Jul 30
Q-Steer uses rollout-time action-value steering to improve molecular policy optimization with oracle-limited feedback
arXiv cs.LG
source ↗
Jul 30
Steering instruction hierarchies at inference time
arXiv cs.CL
source ↗
Jul 30
Flow map learning via nongradient vector flow
arXiv cs.LG
source ↗
Jul 30
Weak-to-Strong on-policy distillation improves alignment when no larger teacher exists
arXiv cs.LG
source ↗
Jul 30
Evaluating prompt scope and demonstration similarity in local LLM machine translation
arXiv cs.CL
source ↗
Jul 30
(Im)Paired programming: coding agents improve productivity but harm understanding
arXiv cs.CL
source ↗
Jul 30
Robostreet Flow: a battery-electric 6x4 tractor with a four-truck convoy architecture to minimize cost per ton-mile on high-volume point-to-point freight
arXiv cs.CL
source ↗
Jul 29
MIT-developed PhysioNet evolves into a global data-sharing standard for medical research.
MIT News (AI)
source ↗
Jul 29
K-Search translates CUDA optimization knowledge into architecture-native MLX strategies for Apple Silicon
BAIR (Berkeley)
source ↗
Jul 29
SpecPrefetch: Parameter-efficient expert prefetching for sparse MoE foundation models
arXiv cs.AI
source ↗
Jul 29
VisualPatchWorld: Code world models as latent structured representations for planning
arXiv cs.CL
source ↗
Jul 29
Human preference aligned tabular similarity for similarity search in business systems; arXiv:2607.24880 highlights limitations of standard metrics and argues for embedding trustworthiness in human-aligned rankings
arXiv cs.LG
source ↗
Jul 29
Models fake alignment without clear consequences, study suggests
arXiv cs.AI
source ↗
Jul 29
TimeCapsule: a 1.2B LLaMA-style model trained only on Victorian texts (1800–1875) as an epistemologically isolated generative archive.
arXiv cs.CL
source ↗
Jul 29
Evaluation of forced alignment of code-mixed speech: the case of Hindi-English
arXiv cs.CL
source ↗
Jul 29
Unified algorithmic framework for hybrid reinforcement learning in tabular MDPs with shifted transition dynamics
arXiv cs.LG
source ↗
Jul 29
Interpretable column annotation with LLM-symbolized decision process materialization
arXiv cs.CL
source ↗
Jul 29
When shortest isn’t safest: a design science approach to senior-friendly pedestrian routing
arXiv cs.AI
source ↗
Jul 29
Neurai-VN benchmark: standardized machine learning models for multimodal digital phenotyping in mental health classification
arXiv cs.LG
source ↗
Jul 29
LinkedIn presents a unified semantic modeling framework for large-scale job understanding powered by a small language model
arXiv cs.AI
source ↗
Jul 29
GROCLM: a fine-tuned language model for grocery category recommendation in e-commerce
arXiv cs.AI
source ↗
Jul 29
Stable FP4 training via transposition-invariant block quantization
arXiv cs.LG
source ↗
Jul 29
Cross-lingual analysis of entrainment in code-switched speech shows lexical entrainment generalizes across Mandarin-English, Hindi-English, and Spanish-English pairs, while acoustic-prosodic and CSW-style entrainment varies by context.
arXiv cs.CL
source ↗
Jul 29
CAST: Game solvers as turn-level teachers for LLM agents
arXiv cs.CL
source ↗
Jul 29
GAUGE scores financial models without a golden answer using independently built analyst models across 108 directed pairs.
arXiv cs.LG
source ↗
Jul 29
Kernel Forge: an agent harness for LLM-based generation and optimization of CUDA kernels
arXiv cs.AI
source ↗
Jul 29
Score-based stabilization for time-dependent problems in numerical PDE simulation
arXiv cs.LG
source ↗
Jul 29
Phase structure in rotary attention revealed through a bounded spectral framework for semantic continuity and execution-boundary governance
arXiv cs.CL
source ↗
Jul 29
RSMeM: Knowledge-enhanced memory evolution for remote sensing agents with systematic evaluation
arXiv cs.AI
source ↗
Jul 29
HOBA: Hierarchical On-policy Bidding Agents for adaptive online advertising
arXiv cs.AI
source ↗
Jul 29
Right-sizing recommendations for cloud workloads using conformal prediction to select VM sizes in data center operations
arXiv cs.AI
source ↗
Jul 29
Deep Label-Wise Attentive Temporal Convolutional Networks improve medical coding
arXiv cs.CL
source ↗
Jul 29
A scaling law of contextual persistence in human language
arXiv cs.CL
source ↗
Jul 29
Crystalis: Progressive nucleation and semantic annealing for coordinated multi-view visualization generation
arXiv cs.AI
source ↗
Jul 29
Generative Distributionally Robust Optimization (GDRO)
arXiv cs.LG
source ↗
Jul 29
RRS-10K: a benchmark for rare remote sensing image interpretation with 10,738 military-related images
arXiv cs.AI
source ↗
Jul 29
PATHFinder Agent for tailored prenatal care
arXiv cs.AI
source ↗
Jul 29
Mechanisms of width scaling in normalized residual networks: the effective alignment dimension
arXiv cs.LG
source ↗
Jul 29
Multilingual safety gaps in frontier language models: in-context scheming scales with pretraining language coverage, according to Petri auditing findings
arXiv cs.AI
source ↗
Jul 29
LLMs as an alternative to corpora for specialized terminology in English–French translation
arXiv cs.AI
source ↗
Jul 29
IRIS: reusable identity representations from frozen LLMs for entity alignment
arXiv cs.CL
source ↗
Jul 29
Conformal Cascade: distribution-free accuracy guarantees for multi-tier LLM inference
arXiv cs.LG
source ↗
Jul 29
ProcAgent: an agentic framework for procedural task guidance on edge with human-in-the-loop
arXiv cs.AI
source ↗
Jul 29
Physics-Informed CNN-LSTM for street-scale urban flood prediction: reconciling aggregate accuracy and street-level plausibility
arXiv cs.LG
source ↗
Jul 29
Toward a systematic method for identifying language areas
arXiv cs.CL
source ↗
Jul 29
Calibrated partial resets prevent policy collapse in continual reinforcement learning
arXiv cs.LG
source ↗
Jul 29
Interpretable GOHR agents via sparse autoencoders
arXiv cs.LG
source ↗
Jul 29
Semantic space search trajectory networks: constructing STNs in semantic spaces to visualize optimization behavior
arXiv cs.LG
source ↗
Jul 29
Personalization, personas, and forecasting in value alignment
arXiv cs.AI
source ↗
Jul 29
CondPSE: a polynomial-filtered structural encoder with conditional modulation for graphs
arXiv cs.LG
source ↗
Jul 29
LLM-based trace ranking and grouped reward modeling for multilingual numerical claim verification at CheckThat! 2026
arXiv cs.CL
source ↗
Jul 29
Beyond memory: a templated substrate for heterogeneous collaborative knowledge work with LLM agents
arXiv cs.AI
source ↗
Jul 29
Inverse RL helps align AI by imitating humans
arXiv cs.LG
source ↗
Jul 29
TabRank: chain-of-thought distillation for table re-rankers
arXiv cs.CL
source ↗
Jul 29
Reasoning with memory: a temporal granularity-adaptive framework for training-free long video understanding
arXiv cs.AI
source ↗
Jul 29
CaRE: Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models
arXiv cs.AI
source ↗
Jul 29
Rethinking CD: a reproducibility study and extension on the ineffectiveness of contrastive decoding at mitigating object hallucinations in MLLMs
arXiv cs.LG
source ↗
Jul 29
Lantern: Conflict-aware gradient blending for physics-guided diffusion models in calorimeter simulation
arXiv cs.LG
source ↗
Jul 29
RoCo-ACE: Rollout-conditioned online distillation for retention-aware knowledge injection
arXiv cs.AI
source ↗
Jul 29
LLM as forecasting planner: training-free text conditioning for time-series foundation models
arXiv cs.LG
source ↗
Jul 29
Multiclass classification without labels via posterior simplex geometry
arXiv cs.LG
source ↗
Jul 29
Activation source selection in activation steering: where steering signals come from
arXiv cs.CL
source ↗
Jul 29
Steering topology distributions for unified generative design of architected metamaterials
arXiv cs.AI
source ↗
Jul 29
Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2
arXiv cs.LG
source ↗
Jul 29
Linguistic rules alone can serve as effective prompt compressors to reduce inference costs, study finds
arXiv cs.CL
source ↗
Jul 29
Behavior-Driven Explainability
arXiv cs.LG
source ↗
Jul 29
Research report analyzes noise-shaped one-bit coefficients in discrete polynomial Fourier extension; achieves O(N^{-1}) approximation rate for first-order Sigma-Delta quantization
arXiv cs.CL
source ↗
Jul 29
FinAbstain: uncertainty-calibrated multimodal RAG for selective financial forecasting
arXiv cs.LG
source ↗
Jul 29
Endpoint Replay: Compressing the recency buffer in deep reinforcement learning
arXiv cs.LG
source ↗
Jul 29
Atmospheric diffusion-guided spatio-temporal transformer for nuclear radiation forecasting
arXiv cs.AI
source ↗
Jul 28
LeafData: an agentic system that converts user intent into validated JSON configurations for data migration
arXiv cs.AI
source ↗
Jul 28
Transferable latency prediction for fast LLM screening on heterogeneous edge devices
arXiv cs.AI
source ↗
Jul 28
Discrete action space as a prerequisite for GRPO convergence in small-model continuous control
arXiv cs.AI
source ↗
Jul 28
Public-space gesture interaction evaluated over interaction-level success rather than frame-level accuracy; paper analyzes repair records and runtime failure analysis in scenic kiosks and service terminals.
arXiv cs.AI
source ↗
Jul 28
Persistent computational state: a session-centric runtime for generative world models
arXiv cs.AI
source ↗
Jul 28
FrED: External data influence estimation via domain knowledge graph grounding
arXiv cs.AI
source ↗
Jul 28
FlowEvo: self-evolving agents through the co-evolution of workflows and executable skills
arXiv cs.AI
source ↗
Jul 28
Context anxiety causes frontier reasoning models to doubt their solutions despite capability; a study analyzes how misestimation of token difficulty affects performance.
arXiv cs.AI
source ↗
Jul 28
SCOPE and SCION: a benchmark and an auditable reference pipeline for schema induction and fusion from text
arXiv cs.AI
source ↗
Jul 28
FBLayout: Optimizing memory layout for efficient LLM finetuning on mobile GPUs
arXiv cs.AI
source ↗
Jul 28
Spectral Flow Certificates enable depth-aware long-range propagation in graph neural networks
arXiv cs.AI
source ↗
Jul 28
Wavelet phase diffusion for structurally and semantically consistent sim-to-real translation
arXiv cs.AI
source ↗
Jul 28
Do VLMs read or rewrite? Evidence of transcription rewriting in vision-language models using FaithC4 benchmark
arXiv cs.AI
source ↗
Jul 28
HDL emerges as a hard decision layer in transformers, causing abrupt stabilization of answer option rankings during inference
arXiv cs.AI
source ↗
Jul 28
Procedural knowledge is not low-rank: LoRA fails to internalize multi-step procedures
arXiv cs.AI
source ↗
Jul 28
From profiles to steering vectors: Global Sparse Priors and Local Semantic Calibration for personalized text generation
arXiv cs.AI
source ↗
Jul 28
Coupled hierarchical search over topology and execution for agentic workflow synthesis
arXiv cs.AI
source ↗
Jul 28
TILT: Improving compositional generation in diffusion models with a model-intrinsic reward
arXiv cs.AI
source ↗
Jul 28
Trajectory-aware retrieval agents for temporal decision making
arXiv cs.AI
source ↗
Jul 28
Defining AI-native systems: autonomy as revision authority
arXiv cs.AI
source ↗
Jul 28
Do modules stay in their lane? Role drift in compound LLM systems
arXiv cs.AI
source ↗
Jul 28
Securing multimodal AI through internal information decomposition
arXiv cs.AI
source ↗
Jul 28
Monotonic evaluation framework measures whether increases in wildfire risk scores correspond to higher observed operational load
arXiv cs.AI
source ↗
Jul 28
AgentKVShift: Efficient KV cache reuse for agentic memory systems
arXiv cs.AI
source ↗
Jul 28
Household movement detection in mixed-format occupancy data using LLM-based entity resolution
arXiv cs.AI
source ↗
Jul 28
FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow
arXiv cs.LG
source ↗
Jul 28
OrchNAS: orchestrated neural architecture search service for personalised federated edge intelligence
arXiv cs.LG
source ↗
Jul 28
LA-RL: Label-aware self-reflection for reinforcement learning in information extraction
arXiv cs.CL
source ↗
Jul 28
Joint optimization for greedy longest-match tokenization (JOLT) develops a subword vocabulary learned to optimize greedy left-to-right longest-match decoding for WordPiece
arXiv cs.CL
source ↗
Jul 28
CC-AOS: Cost- and horizon-conditioned amortized backward induction for finite-horizon optimal stopping
arXiv cs.LG
source ↗
Jul 28
Integrated deep learning and statistical framework for whole-network gene–environment association with leaf vascular architecture
arXiv cs.LG
source ↗
Jul 28
Memento: memory-guided memetic code-as-policy evolution
arXiv cs.LG
source ↗
Jul 28
QFedPolyp: a communication- and inference-efficient federated learning framework for polyp segmentation
arXiv cs.LG
source ↗
Jul 28
Bharati: morphology-aware tokenizers for classical Indian languages with subword fertility analysis
arXiv cs.CL
source ↗
Jul 28
Activation Oracles learn not to read: concept-specific blind spots in fine-tuned oracles
arXiv cs.CL
source ↗
Jul 28
Beyond Shapley: an influence-based data auditing pipeline for LLM alignment and evaluation
arXiv cs.LG
source ↗
Jul 28
Automated detection of documentation inconsistencies in electronic health records using a two-stage LLM pipeline
arXiv cs.CL
source ↗
Jul 28
Physically verifiable evidence and LLM-based reporting for bearing fault diagnosis
arXiv cs.LG
source ↗
Jul 28
Multimodal domain generalization for depression detection via an attention-based BiLSTM network with domain-adversarial training
arXiv cs.LG
source ↗
Jul 28
Optimizing Transformer neural network for real-time outlier detection on FPGAs
arXiv cs.LG
source ↗
Jul 28
Hierarchical grading in large language models
arXiv cs.LG
source ↗
Jul 28
ADAGE: a language-agnostic pipeline for analogical reasoning evaluation
arXiv cs.CL
source ↗
Jul 28
From hybrid mechanistic–data-driven modeling toward neuro-symbolic AI: what, why, and how
arXiv cs.LG
source ↗
Jul 28
LithoFormer: a robust framework for stratigraphic inference via transformers
arXiv cs.LG
source ↗
Jul 28
IKS-Instruct: a 24,795-pair multilingual dataset for teaching language models Indian Knowledge Systems
arXiv cs.CL
source ↗
Jul 28
CausalGate: causal importance distillation for transformer module pruning
arXiv cs.LG
source ↗
Jul 28
PatiGonit22K: a comprehensive dataset for solving complex Bengali MWPs
arXiv cs.CL
source ↗
Jul 28
DomainPilot: Domain-Level loss-guided two-stage data mixture optimization for efficient language model fine-tuning
arXiv cs.LG
source ↗
Jul 28
Evaluating narrative unlearning with LENS: a level-based evaluation protocol for suppressing disinformation-aligned narrative reproduction in LLMs
arXiv cs.CL
source ↗
Jul 28
IndicTalk: a large-scale multilingual code-mixed conversational corpus for Indic languages
arXiv cs.CL
source ↗
Jul 28
Aligning educational LLMs as Socratic guides via heuristic reinforcement learning; study presents HeuristicEdu for Qwen2.5-7B using SocraticEdu data and GRPO training
arXiv cs.CL
source ↗
Jul 28
AutoThinkSQL enables selective reasoning for text-to-SQL by combining SFT and DPO guidance
arXiv cs.CL
source ↗
Jul 28
MioFFAn: an open-source framework for annotating mathematical expressions for formula formalization with LLM automation capabilities
arXiv cs.CL
source ↗
Jul 28
Multimodal surface sEMG hand gesture recognition using query-based transformers for prosthetic control
arXiv cs.LG
source ↗
Jul 28
LoRA for gender-inclusive rewriting and activation steering for counter-narrative generation achieves 80.00% official score on gender-inclusive rewriting (IHLC system) for LT-EDI 2026 Shared Task
arXiv cs.CL
source ↗
Jul 28
Attention-guided layer selection for contrastive decoding in large language models
arXiv cs.CL
source ↗
Jul 28
Co-evolving graph and text memory for training-free multi-hop question answering
arXiv cs.CL
source ↗
Jul 28
Softmax throws away: mass-aware attention for evidence accumulation
arXiv cs.LG
source ↗
Jul 28
CHiPS: character histograms and positional signals for lightweight authorship attribution in Romanian texts
arXiv cs.CL
source ↗
Jul 28
Kalle Lyytinen discusses the linguistic core of information systems and the developments since his 1980 MIS Quarterly article; interview highlights.
arXiv cs.CL
source ↗
Jul 28
Speech signals complement LLMs for predicting interpersonal attraction in speed dating
arXiv cs.CL
source ↗
Jul 28
Simple language normalization wins: cross-lingual speaker verification for the tidyVoice 2026 challenge
arXiv cs.CL
source ↗
Jul 28
GAND: a resource on gender-ambiguous natural data and contrastive attribution for evaluating gender bias in machine translation
arXiv cs.CL
source ↗
Jul 28
SLA-constrained carbon-aware routing in geo-distributed serverless clouds
arXiv cs.LG
source ↗
Jul 28
Dementia etiology diagnosis via collaborative meta knowledge enhancement
arXiv cs.LG
source ↗
Jul 28
Semalith v1.4, a 184M DeBERTa-v3-base classifier, achieves state-of-the-art prompt-injection detection with 44x fewer parameters than Llama-Guard-3-8B
arXiv cs.LG
source ↗
Jul 28
Predicting rTMS depression therapy outcomes from EEG signals using CNNs
arXiv cs.LG
source ↗
Jul 28
Official conference guidelines yield more consistent results in automated peer review by LLMs.
arXiv cs.CL
source ↗
Jul 28
Progress-conditioned group policy optimization for long-horizon agentic tasks
arXiv cs.LG
source ↗
Jul 28
BERT-based models vs. large language models for low-resource named entity recognition: a comparative study on Marathi
arXiv cs.CL
source ↗
Jul 28
Accessibility plasticity as a principle of adaptive computation for access to computation
arXiv cs.LG
source ↗
Jul 28
Beyond a global norm: personalizing toxicity sensitivity in language models without retraining
arXiv cs.CL
source ↗
Jul 28
Corvus: context optimization and reduction via underlying synchronization for LLM coding agents
arXiv cs.LG
source ↗
Jul 28
LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning
arXiv cs.LG
source ↗
Jul 28
Loss-aware feature-map pruning in convolutional neural networks using multi-armed bandits
arXiv cs.AI
source ↗
Jul 28
SeT-Diff: Towards semantic foundation models for HPC telemetry and time-series
arXiv cs.AI
source ↗
Jul 28
cMoLLM at scale: horizontal scaling laws for mixture-of-LLMs
arXiv cs.AI
source ↗
Jul 28
QFoldAgent: a closed-loop multi-agent framework for 5-residue tetrahedral-lattice protein folding with a design agent proposing sequence-conditioned penalties
arXiv cs.AI
source ↗
Jul 28
PhononBench-MP40 provides a spectrum-resolved benchmark dataset of Materials Project-derived crystals for workflow-defined phonon stability.
arXiv cs.AI
source ↗
Jul 28
Scaffold effect in coding agents: evaluation varies with harness choice across Qwen 3.6 Plus and MiniMax M2.5
arXiv cs.AI
source ↗
Jul 28
Prompt design significantly impacts energy use in on-device LLM inference
arXiv cs.AI
source ↗
Jul 28
HeraSys enables collaborative serving of multiple LLM workflows through fine-grained end-to-end optimization
arXiv cs.AI
source ↗
Jul 28
Evaluating LLM reliability beyond accuracy: how model answers vary with meaning-preserving paraphrases across tasks
arXiv cs.AI
source ↗
Jul 28
Synthetic scenario generation for evaluation of Industry 4.0 agents; extends AssetOpsBench with a Smart Grid Transformer asset class and four IEC-grounded diagnostic tools for health-index prediction, dissolved-gas analysis, and winding-temperature detection.
arXiv cs.AI
source ↗
Jul 28
Execution-Grounded security testing for coding agents in software engineering pipelines
arXiv cs.AI
source ↗
Jul 28
SCEPTER converts clinical case descriptions into evidence-based recommendations using multi-objective evidence reasoning
arXiv cs.AI
source ↗
Jul 28
Schema-Aware Localisation (SAL): live schema grounding and hallucination validation for Oracle NL2SQL
arXiv cs.AI
source ↗
Jul 28
Codifying the Judge: scalable evaluation via program distillation
arXiv cs.AI
source ↗
Jul 28
Memory-Induced Inference-Time Adaptation for continual learning with small language models
arXiv cs.AI
source ↗
Jul 28
Source-aware reranking for retrieval-augmented generation using a reliability prior approach
arXiv cs.AI
source ↗
Jul 28
Multi-Objective structured pruning of LLMs for latency and model size optimization
arXiv cs.AI
source ↗
Jul 28
DeepLens Diagnosis Agent uses a five-stage reasoning pipeline to enable a small medical reasoning model to compete with frontier LLMs.
arXiv cs.AI
source ↗
Jul 28
Concept-based visual counterfactual explanations with diffusion models
arXiv cs.AI
source ↗
Jul 28
Reference feature atlases for mechanistic auditing of language models
arXiv cs.AI
source ↗
Jul 28
SCAIR: Schema-conditioned agentic iterative reasoning for enterprise knowledge graphs
arXiv cs.AI
source ↗
Jul 28
DSTFView: multi-view workload forecasting for cloud-edge platforms using dual-input spatio-temporal-frequency modeling
arXiv cs.AI
source ↗
Jul 28
MedLoCoMo: a long-context, multi-session medical dialogue benchmark for large language models
arXiv cs.AI
source ↗
Jul 27
Self-poisoning in adaptive out-of-distribution detection: a sharp-threshold theory and certified label-free calibration
arXiv cs.LG
source ↗
Jul 27
Benchmarking fine-tuning and retrieval strategies for a multimodal language model on the NRC Reactor Operator licensing examination
arXiv cs.CL
source ↗
Jul 27
Molt: a PyTorch-native training framework for agentic reinforcement learning reduces researcher cost per iteration
arXiv cs.LG
source ↗
Jul 27
Quasi-Monte Carlo initialization improves meta-reinforcement learning training convergence in benchmark environments; QMC meta-priors outperform SB3 defaults on unseen continuous control tasks.
arXiv cs.LG
source ↗
Jul 27
Shallower ReLU network representations via exact linear algebra
arXiv cs.LG
source ↗
Jul 27
Probing latent Colombian identity inferences in Qwen2.5-7B with natural language autoencoders
arXiv cs.CL
source ↗
Jul 27
Multi-horizon latent consistency as geometry: increasing lambda affects empirical expansion proxy L20 and horizon-20 prediction error on Moving-MNIST
arXiv cs.LG
source ↗
Jul 27
Quadratic model can be surprisingly predictive for optimization in large-language models with 150M parameters
arXiv cs.LG
source ↗
Jul 27
A consensus-based framework for relative preference evaluation of large language models
arXiv cs.CL
source ↗
Jul 27
Learning what matters: supervising sparse attention routing with causal evidence sets
arXiv cs.LG
source ↗
Jul 27
Reinforcement learning mitigates task conflicts in LLM model merging; analysis examines training paradigms' impact
arXiv cs.CL
source ↗
Jul 27
Encoding invisible causation for bridge diagnostic agents; triple-guided retrieval-augmented fine-tuning with QLoRA
arXiv cs.LG
source ↗
Jul 27
Khondo: a multimodal benchmark for document packet splitting of Bangla forms
arXiv cs.CL
source ↗
Jul 27
Bounding the causal impact of ML-assisted decision-making via counterfactual correctness
arXiv cs.LG
source ↗
Jul 27
A drift-stable quantum federated learning for intelligent services
arXiv cs.LG
source ↗
Jul 27
Adjustment speed as a safety constraint for nonstationary reinforcement learning
arXiv cs.LG
source ↗
Jul 27
FSE: Continual learning for named entity recognition by fast-slow experts
arXiv cs.CL
source ↗
Jul 27
Neural Atom Prevalence extends atom prevalence to neural networks for structured node-level model selection in feedforward architectures.
arXiv cs.LG
source ↗
Jul 27
Analyzing self-harm representations in language models: a cross-architecture study
arXiv cs.CL
source ↗
Jul 27
From isolated tasks to structured capabilities: a multilayer taxonomy for large language models
arXiv cs.CL
source ↗
Jul 27
Physiological signals as a forensic modality for talking-face deepfake detection
arXiv cs.LG
source ↗
Jul 27
Spanish version of the Large Language Models Dependency Scale (LLM-D12-SP) developed and validated
arXiv cs.CL
source ↗
Jul 27
Measuring the dependency gap: diagnosing inter-column fidelity in tabular generative models
arXiv cs.LG
source ↗
Jul 27
Cloud-Native Evaluation-as-a-Service: a six-microservice architecture for scalable AI monitoring with conformal guarantees
arXiv cs.LG
source ↗
Jul 27
Dynamic Commonsense Coordination for empathetic response generation
arXiv cs.CL
source ↗
Jul 27
Agentic evaluation of copyright law compliance with Copyright-Bench assessing LLM agents’ compliance with copyright law
arXiv cs.CL
source ↗
Jul 27
Data quality over capacity: baking documents into LoRA adapters for closed-book QA
arXiv cs.CL
source ↗
Jul 27
Smart predict-then-robustly-optimize: account for prediction shifts due to disturbance in the covariate feature space
arXiv cs.LG
source ↗
Jul 27
RED-PIM: Reducing data movement for transformers using processing-in-memory
arXiv cs.LG
source ↗
Jul 27
Ground Truth First: a longitudinal evaluation instrument for agent memory and the tenure crossover in memory-architecture rankings
arXiv cs.CL
source ↗
Jul 27
Humanly: A configurable and traceable environment for human-AI collaborative writing
arXiv cs.CL
source ↗
Jul 27
MEUSLI: a multilingual projector linking Whisper encoder with open-source multilingual ASR and beyond
arXiv cs.CL
source ↗
Jul 27
Scaling native multimodal pre-training from scratch
arXiv cs.CL
source ↗
Jul 27
Introduction to Bayesian and frequentist simulation-based inference with machine learning
arXiv cs.LG
source ↗
Jul 27
Improving faithfulness of podcasts from documents
arXiv cs.CL
source ↗
Jul 27
Restoration of historical documents via retrieval-augmented large language models leverages external knowledge to recover illegible entities
arXiv cs.CL
source ↗
Jul 27
J-CoT: Chain-of-Thought in J-Space; explores latent-reasoning through recurrent propagation of continuous hidden states
arXiv cs.CL
source ↗
Jul 27
Depth scalability limitations in logic gate networks; optimization collapse and topology-induced constraints persist beyond training improvements
arXiv cs.LG
source ↗
Jul 27
Toward user-conditioned evaluation of personal LLM agents under temporal interventions
arXiv cs.LG
source ↗
Jul 27
Analyzing toxic behavior and its impact on the Mastodon community
arXiv cs.CL
source ↗
Jul 27
Parameter-free adaptive sparse attention via compression-based content selection
arXiv cs.LG
source ↗
Jul 27
Goal-agnostic joint-embedding predictive control framework for partial differential equations using offline-trained ViT encoder and action-conditioned latent dynamics with an MPPI controller
arXiv cs.LG
source ↗
Jul 27
DWT-Fusion: a signal-based framework for training-free LLM-generated text detection
arXiv cs.CL
source ↗
Jul 27
MoE$^2$-LoRA: when MoE models meet MoE-style low-rank adaptation
arXiv cs.CL
source ↗
Jul 27
Evaluation design conditions the expert-vs-auto MeSH gap; a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark
arXiv cs.CL
source ↗
Jul 27
Adversarial style optimization: enhancing VLM jailbreaks by GRPO-based stylistic triggers optimization
arXiv cs.CL
source ↗
Jul 27
Physically constrained federated additive models for O-RAN SLA-risk prediction
arXiv cs.LG
source ↗
Jul 27
MotifRole-Diff: Risk-optimal role-aware corruption for masked molecular graph diffusion
arXiv cs.LG
source ↗
Jul 27
CARNet: cycle-conditioned core aggregation and redistribution for multivariate time series forecasting
arXiv cs.LG
source ↗
Jul 26
Teaching LLMs to update beliefs for efficient long-horizon interaction
BAIR (Berkeley)
source ↗
Jul 25
What is good? Extracting and testing implicit theories of literary quality from LLM reasoning traces
arXiv cs.CL
source ↗
Jul 25
Naver-News-KO: a Korean news summarization dataset for open-source fine-tuning of summarization models
arXiv cs.CL
source ↗
Jul 25
More Is Not More: what matters for diversity in LLM opinions?
arXiv cs.CL
source ↗
Jul 25
Answer-then-Edit: Reasoning skeleton editing for anti-distillation with preserved utility
arXiv cs.CL
source ↗
Jul 25
AsymVerify achieves 2nd place on SemEval-2026 Task 6 with a confidence-gated verification approach for political evasion detection
arXiv cs.CL
source ↗
Jul 25
LLM-INSTRUCT wins UZH Shared Task 2026 on paragraph-level argument mining with constraint-aware retrieval and selective debate
arXiv cs.CL
source ↗
Jul 25
Natural language should not fully replace formal languages, argues a position paper highlighting underspecification in open-ended contexts
arXiv cs.CL
source ↗
Jul 25
A three-stage neural-symbolic pipeline using an ensemble of DeBERTa-v3-base and XLM-RoBERTa-base with a Linguistically-Informed Mediator for gaming toxicity detection.
arXiv cs.CL
source ↗
Jul 25
TopoGuard: graph theory based defenses against split-knowledge attacks on RAG
arXiv cs.CL
source ↗
Jul 25
Routing Subspaces: Auditing evaluation-to-deployment mismatch in fine-tuned language models
arXiv cs.CL
source ↗
Jul 25
Open-source text LLM watermarks fail to withstand model merging, study finds
arXiv cs.CL
source ↗
Jul 25
Human-in-the-loop LLM framework improves detection of cutaneous immune-related adverse events in clinical notes; study reports higher accuracy and agreement vs manual review
arXiv cs.CL
source ↗
Jul 25
MoE routing follows a Huffman code pattern, as the Frequency-Diversity Law shows that state-of-the-art models act as information-theoretic engines.
arXiv cs.CL
source ↗
Jul 25
Skill-contracted agents for evidence-aware materials literature analysis
arXiv cs.CL
source ↗
Jul 25
SCoPE: Shift-aware speaker-conditioned priors for emotion recognition in conversations
arXiv cs.CL
source ↗
Jul 25
Knowledge injection exists in MoE; exploring expert-aware contrast decoding in MoE for mitigating LLMs' hallucinations
arXiv cs.CL
source ↗
Jul 25
Moir: Let the model direct its own story for robust cross-domain knowledge editing
arXiv cs.CL
source ↗
Jul 25
Distinguishing artificial from authentic: evaluating LLMs for detecting LLM-generated content
arXiv cs.CL
source ↗
Jul 25
GLAN-QnA-KR: a 303,581-row Korean instruction-QA corpus produced by seedless taxonomy-driven GLAN synthesis using Microsoft Phi-3.5-MoE-instruct; open and redistributable under OpenRAIL
arXiv cs.CL
source ↗
Jul 25
Preference tuning as spectral update reorganization
arXiv cs.CL
source ↗
Jul 25
Confidently deceptive: how confidence amplifies the risk of LLM deception
arXiv cs.CL
source ↗
Jul 25
Break Through the Compression Bottleneck: From Theory to Practice
arXiv cs.CL
source ↗
Jul 25
The Storyteller in the Model: Narrative pattern inheritance, escalation dynamics, and alignment governance in LLMs
arXiv cs.CL
source ↗
Jul 25
Belief propagation in LLM world models: measuring strategic information bias with prediction markets
arXiv cs.CL
source ↗
Jul 24
HypNO: a graph-based neural operator for scalar hyperbolic conservation laws uses physics-informed message passing on a space-time graph of finite-volume cells
arXiv cs.LG
source ↗
Jul 24
Conflict resolution under degraded surveillance in air corridors using multi-agent reinforcement learning
arXiv cs.LG
source ↗
Jul 24
AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics
arXiv cs.AI
source ↗
Jul 24
Adaptive depth in looped transformers: diagnosing learned halting gates and trajectory readouts
arXiv cs.LG
source ↗
Jul 24
Uncertainty-aware trust estimation for multi-LLM systems via structured expert judgement
arXiv cs.LG
source ↗
Jul 24
ClickGuard: Detecting and spoiling clickbait news with informativeness measures and large language models
arXiv cs.AI
source ↗
Jul 24
DC-Leap: training-free acceleration of dLLMs via draft-guided contiguous leaping decoding
arXiv cs.AI
source ↗
Jul 24
The Devil is in the spectrum: mitigating representation collapse in LLMs via topologically regularized side-path
arXiv cs.AI
source ↗
Jul 24
Incomplete prompt jailbreaks in large language models reveal vulnerabilities to harmful continuations in open-weight models
arXiv cs.AI
source ↗
Jul 24
Graph neural network approach to zero-shot digital twins
arXiv cs.LG
source ↗
Jul 24
Codec-Gauge learns equalization transforms for transformer KV-cache compression to improve fidelity
arXiv cs.LG
source ↗
Jul 24
PersonaTrail benchmarks personalized web agents through browsing histories to evaluate how agents infer context from user browsing data
arXiv cs.AI
source ↗
Jul 24
OPTScientist: multi-agent discovery of typed optimizer programs for transformer pretraining
arXiv cs.AI
source ↗
Jul 24
DataPrep-Bench benchmarks LLMs as training data preparators; evaluates data construction and data quality evaluation end to end
arXiv cs.LG
source ↗
Jul 24
Generative Bayesian filtering for state estimation
arXiv cs.LG
source ↗
Jul 24
Leveraging biokinetic knowledge priors for data-scarce bioprocess modeling
arXiv cs.LG
source ↗
Jul 24
When RLVR shrinks the reasoning boundary: diagnosing pass@k inversion
arXiv cs.LG
source ↗
Jul 24
Lie typology, depth, and sparsity affect deception detection in LLM outputs
arXiv cs.AI
source ↗
Jul 24
SOAP, Muon, and beyond: pushing LLM pretraining scales
arXiv cs.LG
source ↗
Jul 24
PhantomFill: when the form demands an answer, language models invent one
arXiv cs.LG
source ↗
Jul 24
Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating
arXiv cs.AI
source ↗
Jul 24
SevDiff: Severity-conditioned diffusion for long-tail conflict trajectory generation
arXiv cs.LG
source ↗
Jul 24
Muon optimizer reaches grokking threshold on modular arithmetic faster than AdamW; ablation shows speedup from Newton-Schulz orthogonalization
arXiv cs.LG
source ↗
Jul 24
Stochastic sampling is epistemically shallow; the dimensionality gap between temperature variation and model diversity in LLMs
arXiv cs.AI
source ↗
Jul 24
DecodeShare identifies a shared low-dimensional subspace in decode-time hidden states and tests its causal role by removing it during decoding
arXiv cs.AI
source ↗
Jul 24
Tractable hierarchical control of autoregressive language models
arXiv cs.AI
source ↗
Jul 24
Scaling closed-loop feature channel configuration with LLMs
arXiv cs.LG
source ↗
Jul 24
Benchmarking the personalization capabilities of large language models
arXiv cs.AI
source ↗
Jul 24
Grounding investor views: neural predicates in the Black-Litterman model
arXiv cs.LG
source ↗
Jul 24
Semi-supervised text-attributed graph distillation
arXiv cs.AI
source ↗
Jul 24
Directional hallucinations: ideological drift in news-grounded LLM question answering
arXiv cs.AI
source ↗
Jul 24
Enabling scalable topology inference in distribution systems via constrained multi-source inference
arXiv cs.AI
source ↗
Jul 24
SenCos-GEM: SENet-Calibrated and Law-of-Cosines-Constrained geometry-enhanced molecular representation for property prediction
arXiv cs.LG
source ↗
Jul 24
PlanE: Meta planning of data, tuning, and inference for extractive-based LLMs
arXiv cs.AI
source ↗
Jul 24
Marking the wrong symptoms: evaluating LLM watermarks in medical texts
arXiv cs.AI
source ↗
Jul 24
ReliableTableQA evaluates the amount of supervision needed for reliability annotation of tabular QA results
arXiv cs.LG
source ↗
Jul 24
Expectation alignment of language models for real-world user expectations
arXiv cs.AI
source ↗
Jul 24
Beyond SBDD: geometric deep learning in polypharmacology and multi-target drug design
arXiv cs.LG
source ↗
Jul 24
Decision-aware machine learning framework proposed to improve allocation of essential medicines; arXiv paper introduces multi-task approach
arXiv cs.LG
source ↗
Jul 24
JAXBench: Benchmarking autonomous TPU kernel optimization
arXiv cs.AI
source ↗
Jul 24
Proactive test-driven AI development is proposed to replace reactive patching of models based on user feedback.
arXiv cs.LG
source ↗
Jul 24
Optimal noise allocation for diffusion training in the convex regime
arXiv cs.LG
source ↗
Jul 24
Multimodal CoLRAG-TF uses four-axis fusion—dense text embeddings, BM25, knowledge-graph triple filtering, and image similarity—to enable robust retrieval over complex PDFs.
arXiv cs.LG
source ↗
Jul 24
SonicSampler: unified tile-aware kernels for LLM sampling and speculative verification
arXiv cs.AI
source ↗
Jul 24
InferenceBench: a benchmark for open-ended LLM inference optimization by AI agents
arXiv cs.AI
source ↗
Jul 24
CLOE: Christoffel Loss Autoencoder for anomaly detection
arXiv cs.LG
source ↗
Jul 24
Active SAE feature planes carry more holonomy in Gemma 2 2B, preregistered reversal findings
arXiv cs.LG
source ↗
Jul 24
Automating nuclear plant operations through collaboration and stakeholder relationships, says Lauren Fortier
MIT News (AI)
source ↗
Jul 23
MIT projects selected for funding under U.S. Department of Energy’s Genesis Mission
MIT News (AI)
source ↗
Jul 23
BatchDAG: LLM-planned execution graphs for scalable ad-hoc analysis over enterprise data
arXiv cs.AI
source ↗
Jul 23
SysAdmin: measuring instrumental power-seeking in frontier AI
arXiv cs.AI
source ↗
Jul 23
From Agent Failure Paths to Quantified Residual Risk: A Compositional Framework for Resilient Agentic AI
arXiv cs.AI
source ↗
Jul 23
Cross-dialect generalization without retraining: benchmarks and evaluation of schema-derived constrained decoding for MLIR
arXiv cs.AI
source ↗
Jul 23
Semantic cooperative games for contribution attribution in LLM-based multi-agent systems
arXiv cs.AI
source ↗
Jul 23
Probabilistic concept-aware steering for trustworthy LLM inference
arXiv cs.AI
source ↗
Jul 23
PEARL: solver-in-the-loop interactive optimization modeling from natural language
arXiv cs.AI
source ↗
Jul 23
Phionyx introduces a deterministic AI runtime architecture with structured state management and governance-first state evolution; LLM outputs are treated as noisy sensor measurements rather than decisions.
arXiv cs.AI
source ↗
Jul 23
SAAG: Structured Agent Assessment and Grounding
arXiv cs.AI
source ↗
Jul 23
FindStatBench: evaluating large language models on combinatorial code synthesis
arXiv cs.AI
source ↗
Jul 23
MUX: continuous reasoning via multiplexed tokens
arXiv cs.AI
source ↗
Jul 23
Integro-differential equations in angular stabilization of drone motion by distributed feedback control
arXiv cs.AI
source ↗
Jul 23
Structured synthetic reasoning data improves arithmetic fine-tuning for small language models, using a 21,250-example GPT-5-mini-generated corpus derived from GSM8K
arXiv cs.AI
source ↗
Jul 23
Trajectory-aware clinical risk prediction via severity-grounded knowledge graphs and retrieval-augmented generation
arXiv cs.AI
source ↗
Jul 23
Wisdom of LLM crowds: aggregation and contamination in language model ensembles
arXiv cs.AI
source ↗
Jul 23
Calibrated selective fact-checking via evidence chain evaluation.
arXiv cs.AI
source ↗
Jul 23
State compression in two-agent LLM relays preserves constraints in a closed-world travel-planning relay, a study of hand-off compression effects
arXiv cs.AI
source ↗
Jul 23
MILP-Evo enables closed-loop, fully automatic design of MILP solvers
arXiv cs.AI
source ↗
Jul 23
ProbSPARQL: querying knowledge graphs with multi-dimensional, uncertain numeric data
arXiv cs.AI
source ↗
Jul 23
ToolDNS enables semantic tool discovery over DNS to scale autonomous AI agent tool use; arXiv paper proposes retrofitting discovery onto DNS to handle millions of tools
arXiv cs.AI
source ↗
Jul 23
When JSON is not enough: semantic reliability of schema-constrained LLM ordering agents
arXiv cs.AI
source ↗
Jul 23
Latency-aware LLM query routing for dynamic workloads improves by considering generation latency alongside accuracy and cost
arXiv cs.AI
source ↗
Jul 23
AI/ML deepfake research misaligned with AI-generated non-consensual intimate imagery, according to landscape analysis of highly cited works
arXiv cs.AI
source ↗
Jul 23
Fence: Specialized SLM guardrails for LLM applications
arXiv cs.AI
source ↗
Jul 23
Bayesian wind tunnels for model selection
arXiv cs.LG
source ↗
Jul 23
TalentCLEF at CLEF2026: second edition evaluates NLP systems for fair, multilingual talent and job title intelligence in human capital management
arXiv cs.CL
source ↗
Jul 23
Rubric-oriented document set selection and ranking for retrieval beyond relevance-centric methods
arXiv cs.CL
source ↗
Jul 23
Cross-subject semantic decoding with shared-space alignment for generalized neural representation learning
arXiv cs.LG
source ↗
Jul 23
Reproducing Recurrent Transformers: The CoTFormer
arXiv cs.LG
source ↗
Jul 23
Language-specific versus cross-lingual knowledge graphs for implicit aspect identification in Arabic: a comparative study of reasoning and adaptation strategies
arXiv cs.CL
source ↗
Jul 23
Memory Merge DQN: sensitivity weighted target updates for stable value learning
arXiv cs.LG
source ↗
Jul 23
LAARA: Layer-Aware Adaptive Rank Allocation for parameter-efficient fine-tuning
arXiv cs.LG
source ↗
Jul 23
Task competence is not instruction following: evaluating instruction-conflicting behavior in small language models
arXiv cs.CL
source ↗
Jul 23
Decodable but Not Detectable: A Leakage Fingerprint for Near-OOD Benchmarks reveals a benchmark leak where removing the OOD class and retraining 35 models raises AUROC from 0.326 to 0.911.
arXiv cs.LG
source ↗
Jul 23
Neural operator surrogates for two-dimensional neutron flux estimation.
arXiv cs.LG
source ↗
Jul 23
Scaling laws for hypernetwork-based knowledge injection in large language models
arXiv cs.CL
source ↗
Jul 23
Lightweight system for person–place relation extraction in historical newspapers using dependency graphs and proximity features
arXiv cs.CL
source ↗
Jul 23
Multi-dimensional evaluation of explainability in media bias detection using the BABE dataset
arXiv cs.CL
source ↗
Jul 23
STN-TGAT: Top-K portfolio construction via prior-guided graph attention with learnable soft-threshold sparsification
arXiv cs.LG
source ↗
Jul 23
Sentence Splitter uncovers latent factual structure of sentences with a T5-based self-supervised framework for identifying semantic head-tail boundaries
arXiv cs.CL
source ↗
Jul 23
Predicting groundwater arsenic concentrations using graph neural networks
arXiv cs.LG
source ↗
Jul 23
VizRAG enhances retrieval-augmented generation with hypergraph visualization to organize complex n-ary facts among entities beyond binary relationships
arXiv cs.CL
source ↗
Jul 23
Recovering clinical utility under differential privacy: empirical validation of adaptive federated aggregation on heterogeneous cardiovascular datasets
arXiv cs.LG
source ↗
Jul 23
Predictive single cell foundation model for gene regulation and aging with privacy-preserving tabular learning
arXiv cs.LG
source ↗
Jul 23
Stateful guardrails for multi-turn LLM systems address conversational risk accumulation by tracking session-layer trajectory signals to detect gradual intent drift, fragmented instruction assembly, and repeated disclosures.
arXiv cs.CL
source ↗
Jul 23
Air Quality Arena: a large-scale multi-region ground monitoring dataset and benchmark for air quality forecasting with time-series foundation models
arXiv cs.LG
source ↗
Jul 23
Explainability challenges in continual learning for time series forecasting with Experience Replay strategies
arXiv cs.LG
source ↗
Jul 23
Reasoning narrows the move: diversity collapse in LLM game play
arXiv cs.CL
source ↗
Jul 23
Orthogonalized read improves noisy associative recall in mLSTM memory by reconditioning the learning problem during training plateau (arXiv:2607.19390)
arXiv cs.LG
source ↗
Jul 23
Native multi-dimensional subquadratic operators via input dependent long convolutions
arXiv cs.LG
source ↗
Jul 23
SUM: Unified Geometric Surgery on Spatio-Temporal Adaptation Vectors for Federated Class Incremental Learning
arXiv cs.LG
source ↗
Jul 23
On the computational complexity of structural generalization
arXiv cs.CL
source ↗
Jul 23
Pipeline choices dominate autointerpretability score variance in sparse autoencoder evaluations
arXiv cs.LG
source ↗
Jul 23
Embedding-based measurement of data diversity for NLP; emb-diversity tool assesses diversity using embeddings
arXiv cs.CL
source ↗
Jul 23
Reliability-aware distillation can harm some samples in low-resource Bangla summarization; standard knowledge distillation yields negligible ROUGE-L gains on BanSum Bangla
arXiv cs.CL
source ↗
Jul 23
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi.
arXiv cs.CL
source ↗
Jul 23
Adaptive capitulation: a structural failure mode of LLM responses in vulnerability contexts
arXiv cs.CL
source ↗
Jul 23
Reference-free evaluation of reasoning in open-ended question answering reveals a reasoning-based auditing framework that decomposes reasoning traces into segments and labels premise-target relations with natural language inference.
arXiv cs.CL
source ↗
Jul 23
When does consensus beat voting? A critical analysis of statistical label fusion in medical image segmentation
arXiv cs.LG
source ↗
Jul 23
From trajectories to prefixes: reusing teacher trajectories via replayed prefixes and online continuation
arXiv cs.LG
source ↗
Jul 23
CruiseBench: a real-flight-aligned N-CMAPSS benchmark for engine RUL prediction
arXiv cs.LG
source ↗
Jul 23
Reward-aware population scaling of evolutionary strategies in LLM fine-tuning
arXiv cs.LG
source ↗
Jul 23
Scale-aware learning of chaotic dynamics on unstructured meshes via binned spectral losses
arXiv cs.LG
source ↗
Jul 23
NMR elucidation treated as an agentic search problem; an autonomous agent using a frozen LLM matches graduate-level chemistry students in structure determination
arXiv cs.LG
source ↗
Jul 23
Tiny_schiller: a drop-in German drama corpus for small language models
arXiv cs.CL
source ↗
Jul 23
TriAgent: Divergence-aware multi-agent committees for cost-efficient financial sentiment analysis
arXiv cs.CL
source ↗
Jul 23
Learning the Arabic dialect continuum as a continuous space: a regression approach to speaker origin prediction
arXiv cs.CL
source ↗
Jul 23
Mitigating scaffolding collapse in Socratic tutors via representation alignment
arXiv cs.AI
source ↗
Jul 23
Beyond Tracking or Shortcut: Composition-bounded predictive states in poker autoregressive models
arXiv cs.AI
source ↗
Jul 23
Geometry-Guided Constraint Learning for LLM safety classification achieves near-perfect per-category accuracy with sparse autoencoder feature extraction, reducing constraint counts from K=4-25 to K=2 for most categories
arXiv cs.AI
source ↗
Jul 23
Benchmarking confidential GPU inference on NVIDIA H100 under Intel TDX
arXiv cs.AI
source ↗
Jul 23
NEXUS: Structured runtime safety for tool-using LLM agents
arXiv cs.AI
source ↗
Jul 23
FormulaSPIN: Self-play fine-tuning for natural language to spreadsheet formula generation
arXiv cs.AI
source ↗
Jul 23
LLMs update memory through shared latent structures rather than isolated facts, supporting the lifted representation hypothesis
arXiv cs.AI
source ↗
Jul 23
FORCE-Bench: a benchmark, dataset, and evaluation harness for agentic AI in enterprise finance
arXiv cs.AI
source ↗
Jul 23
CrackedPDFs: a controlled benchmark for hidden prompt injection in PDFs
arXiv cs.AI
source ↗
Jul 23
OpenEvoShield: Dual non-stationary continual defense for open-world multi-agent system attacks
arXiv cs.AI
source ↗
Jul 23
ITPEval benchmarks automated translation of formal proofs across Lean 4, Rocq, Isabelle, and HOL Light
arXiv cs.AI
source ↗
Jul 23
HyGRL: Adaptive Hybrid Graph Reasoning for Multi-Entity Questions
arXiv cs.AI
source ↗
Jul 23
AdaRoPE: not all attention heads should rotate and scale equally
arXiv cs.AI
source ↗
Jul 23
Spectral-LSH: sub-quadratic prompt compression via Krylov-projected locality-sensitive hashing
arXiv cs.AI
source ↗
Jul 23
Statistically grounded sparse-feature interventions for activation-space control in large language models
arXiv cs.AI
source ↗
Jul 23
Stochastic primal-dual decoding for multiobjective generative recommender systems
arXiv cs.AI
source ↗
Jul 23
GraphContainer provides a unified platform for comparing and debugging Graph RAG methods
arXiv cs.AI
source ↗
Jul 23
Hybrid LSTM-Graph neural framework for robust financial fraud detection and adversarial resilience
arXiv cs.AI
source ↗
Jul 23
Euclean: automated geometry problem formalization with unified verification in Lean
arXiv cs.AI
source ↗
Jul 23
LISA: Linear-Indexed Sparse Attention for efficient long-context reasoning
arXiv cs.AI
source ↗
Jul 23
FineServe: a fine-grained dataset and characterization of global LLM serving workloads
arXiv cs.AI
source ↗
Jul 22
SymptomAI: a conversational AI agent for everyday symptom assessment
Google Research
source ↗
Jul 22
Dimitri Bertsekas, professor emeritus and influential computer scientist, dies at 83.
MIT News (AI)
source ↗
Jul 22
Stochastic Meta-Unlearning: Bridging language backbone and multimodal unlearning
arXiv cs.CL
source ↗
Jul 22
Beyond single-dimensional compression: exploring compound sparsity to delay performance degradation in large language models
arXiv cs.LG
source ↗
Jul 22
On the limits of support-preserving alignment and bounded filtering; study examines whether alignment plus safety filters can drive harmful behavior to zero in large language models
arXiv cs.LG
source ↗
Jul 22
Dual Attention Residuals in transformers enable two complementary residual pathways for historical retrieval and multi-stream trajectories; DAR integrates both mechanisms
arXiv cs.CL
source ↗
Jul 22
European Multilingual Evaluation Dataset localized into 11 languages via the EMT network project; DGT and EMT collaborate on translating MMLU for LLM benchmarking
arXiv cs.CL
source ↗
Jul 22
Structured Output collapses answer diversity across 44 language models
arXiv cs.CL
source ↗
Jul 22
FedCC: a low-resource federated adaptation of foundation models for robust corpus callosum localization in fetal ultrasound images
arXiv cs.LG
source ↗
Jul 22
RF-Agent: A practical framework for building language agents for RFIC design
arXiv cs.CL
source ↗
Jul 22
PathReportEval provides a systematic benchmark for pathology report generation
arXiv cs.CL
source ↗
Jul 22
Convolution for large language models; study finds depthwise convolutions provide local inductive bias in Qwen3 Transformer blocks
arXiv cs.CL
source ↗
Jul 22
Interpreting how instruction-tuned transformers encode discourse relations, focusing on causation and antithesis, in next-token prediction tasks
arXiv cs.CL
source ↗
Jul 22
BearingNAS enables in-sensor fault diagnosis for bearings on a laptop with hardware-aware neural architecture search and extreme micro-budget constraints
arXiv cs.LG
source ↗
Jul 22
A domain-conditional position offset reduces the cold-start penalty for autoregressive language models by adding a learned vector to the first token embeddings while freezing model weights.
arXiv cs.LG
source ↗
Jul 22
Dual-domain fused LSTM model for efficient time-dependent reliability analysis
arXiv cs.LG
source ↗
Jul 22
Relay-Bench evaluates LLMs on multi-domain reasoning chains; GPT-5.5 (xHigh) scores 43.3% on composite problems across domains
arXiv cs.CL
source ↗
Jul 22
Reliability declines as model scale increases due to an auto-regressive risk regime that amplifies mistakes
arXiv cs.LG
source ↗
Jul 22
Gradient-energy guided block-wise perturbations for sharpness-aware minimization
arXiv cs.LG
source ↗
Jul 22
Spectral evidence bundling for selective reliability estimation in time-series classification
arXiv cs.LG
source ↗
Jul 22
From a multilingual streaming ASR backbone to Kenyan-language systems: data-centric adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin
arXiv cs.CL
source ↗
Jul 22
E-SpecFormer: an edge-efficient transformer for end-to-end RF spectrum monitoring with LiTAN attention and four scalable variants
arXiv cs.LG
source ↗
Jul 22
FALCON-Discover ranks predictions by discrepancy signals to locate false-confidence regions for calibration
arXiv cs.LG
source ↗
Jul 22
Cost accounting for reactive computational graphs: exhaustive sweeps, sequential mutation, and the backward-locality gap
arXiv cs.LG
source ↗
Jul 22
HPD-Parsing: Hierarchical Parallel Document Parsing
arXiv cs.CL
source ↗
Jul 22
Using fine-tuned LLMs to identify indicators of vulnerability in UK police incident logs
arXiv cs.CL
source ↗
Jul 22
Rationale-Guided Knowledge Distillation for Cross-Lingual Stance Detection
arXiv cs.CL
source ↗
Jul 22
Preference-conditioned multi-objective reinforcement learning for runtime-tunable transit signal priority
arXiv cs.LG
source ↗
Jul 22
AILQA evaluates AI-driven legal question answering for the Indian legal system using embedding and generative models; arXiv researchers assess accuracy and reliability
arXiv cs.CL
source ↗
Jul 22
Reasoning fine-tuning induces persistent latent policy states
arXiv cs.CL
source ↗
Jul 22
Multi-Timescale Latent-Action DRL for joint optimization in edge-cloud networks reduces average end-to-end latency via joint service placement, computational delegation, and power control
arXiv cs.LG
source ↗
Jul 22
Narrative framing affects LLM agent behavior more than persona, as shown by three text-based investigation games with identical structure but different tasks
arXiv cs.CL
source ↗
Jul 22
Fusion Embedding adds audio to a frozen vision-language embedding base to cover text, image, video, and audio in a unified embedding space
arXiv cs.CL
source ↗
Jul 22
Agentic calibration of grey-box simulation models: an LLM-driven alternative
arXiv cs.LG
source ↗
Jul 22
Interactive Training 2: Auditable control plane for live model training
arXiv cs.LG
source ↗
Jul 22
Spatio-temporal prediction of unsteady airfoil aerodynamics using augmented graph neural ordinary differential equations with exogenous controls
arXiv cs.LG
source ↗
Jul 22
Towards principled continual anomaly detection: a systematic framework and benchmark scenarios
arXiv cs.LG
source ↗
Jul 22
Find Before You Fine-Tune: a diagnostic study of small LLMs for cybersecurity QA
arXiv cs.CL
source ↗
Jul 22
Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives
arXiv cs.CL
source ↗
Jul 22
LatentMT: machine translation with latent reasoning
arXiv cs.CL
source ↗
Jul 22
ALAS: Additive Learnable Alpha-Stable Kernels for flexible Bayesian optimization
arXiv cs.LG
source ↗
Jul 22
Self-improving, frozen-gate training (SIFT) enables dynamic document classification via a cheap, CPU-bound pipeline
arXiv cs.CL
source ↗
Jul 22
The information shadow: measuring structural limits on what language models can learn
arXiv cs.LG
source ↗
Jul 22
Uncertainty quantification for AI-driven crash simulation surrogates: Monte Carlo dropout versus deep ensemble on an open-source bumper beam benchmark.
arXiv cs.LG
source ↗
Jul 22
Search-on-Graph-R1: Training an 8B model to search knowledge graphs with reinforcement learning
arXiv cs.CL
source ↗
Jul 21
KernelBench-Verified: LLM-generated kernels often inflate performance; evaluation frameworks must evolve to measure true speedup, not reported gains.
arXiv cs.LG
source ↗
Jul 21
Democratizing AI with small language models: structured benchmarking and parameter-efficient fine-tuning for local deployment
arXiv cs.AI
source ↗
Jul 21
It Takes 8 tokens: weak-to-strong off-policy RL via auxiliary branches
arXiv cs.AI
source ↗
Jul 21
Symbolic augmentation closes a canonical-equivalence blind spot in neural fact-checkers
arXiv cs.AI
source ↗
Jul 21
TalTech submits BeTraC entries: robust summarization of long doctor-patient conversations into SOAP notes using Voxtral models with LoRA fine-tuning and DAPO RL
arXiv cs.CL
source ↗
Jul 21
Literary non-style in LLM-generated text
arXiv cs.CL
source ↗
Jul 21
NOWJ@COLIEE 2026: adaptive pipelines for legal retrieval and reasoning
arXiv cs.CL
source ↗
Jul 21
SEER: supervised learning to control energetic reasoning.
arXiv cs.AI
source ↗
Jul 21
Deterministic replay for AI agent systems
arXiv cs.AI
source ↗
Jul 21
Orthogonal gradient constraints shape noisy-label memorization dynamics
arXiv cs.LG
source ↗
Jul 21
PPO-HSC: An exploratory reinforcement learning framework based on wide-area policy coverage optimization
arXiv cs.AI
source ↗
Jul 21
Are Arithmetic Heuristic Neurons Form-Invariant? A mechanistic analysis of symbols, text, and code in LLMs
arXiv cs.CL
source ↗
Jul 21
When to plan: learning to select between reactive control and deliberative planning
arXiv cs.AI
source ↗
Jul 21
Normalized Rewards for Preference Optimization
arXiv cs.LG
source ↗
Jul 21
Diagnosing correctness probes under self-judgement confounding
arXiv cs.CL
source ↗
Jul 21
A survey on the verification of reinforcement learning policies
arXiv cs.AI
source ↗
Jul 21
Rater state bias in RLHF preference data; an audit framework
arXiv cs.AI
source ↗
Jul 21
A survey on GNN-based link prediction: techniques, applications, and challenges
arXiv cs.AI
source ↗
Jul 21
Auditing question-order effects in large language models with the QQ equality: mechanism characterization and a saturation caveat
arXiv cs.CL
source ↗
Jul 21
Shapley Context Pruning: A cooperative game perspective for context reranking and pruning
arXiv cs.AI
source ↗
Jul 21
Self-Evolving Just-In-Time memory for proactive embodied safety
arXiv cs.LG
source ↗
Jul 21
Operator-aware mixed-precision tolerance calibration for tensor kernels
arXiv cs.LG
source ↗
Jul 21
From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
arXiv cs.LG
source ↗
Jul 21
High-accuracy low-bit KV-cache quantization via local distribution restoration
arXiv cs.LG
source ↗
Jul 21
Learning structural manipulability in gate-level netlists with graph neural networks
arXiv cs.LG
source ↗
Jul 21
Scope3Trace: evidence-based identification and extraction of Scope 3 GHG emissions from sustainability reports
arXiv cs.CL
source ↗
Jul 21
Cascading versus joint modeling for hierarchical offensive language detection
arXiv cs.CL
source ↗
Jul 21
LLM unlearning for cyber defense: a survey on methods, challenges, and emerging threats
arXiv cs.LG
source ↗
Jul 21
RAIL Guard closes the evaluation-to-remediation gap in responsible AI for LLM agents by using a closed-loop pipeline that evaluates outputs across eight dimensions and remediates failures through an evaluate-rewrite-reevaluate loop
arXiv cs.AI
source ↗
Jul 21
RouteCost: a production-inspired multi-stage framework for pre-order shipping cost estimation in e-commerce
arXiv cs.LG
source ↗
Jul 21
RIMS: preference optimization via smoothed multi-pair aggregation for small-scale LLM retrieval-augmented generation
arXiv cs.CL
source ↗
Jul 21
Fully-sensorized smart-eyewear platform enables on-device machine learning via STM32N6 NPU
arXiv cs.LG
source ↗
Jul 21
BACON: Budgeted human calibration for modeling and evaluation with multiple AI judges
arXiv cs.LG
source ↗
Jul 21
Schema-constrained document-level event argument extraction with lightweight LLM fine-tuning
arXiv cs.CL
source ↗
Jul 21
JOR-Bench: Japanese operations research benchmarks for large language models
arXiv cs.CL
source ↗
Jul 21
Learning from synthetic data without model collapse in iterative instruction tuning
arXiv cs.CL
source ↗
Jul 21
Let the data decide: off-policy distillation during continued pre-training analyzes objective-to-capability trade-offs and adaptive objective routing in supervision and performance
arXiv cs.LG
source ↗
Jul 21
Group entropy-controlled policy optimization
arXiv cs.CL
source ↗
Jul 21
Lightweight 1D CNN for affective touch classification in soft plush companions designed and validated; open-source MATLAB framework for compact deep learning models
arXiv cs.AI
source ↗
Jul 21
OpenMHC: accelerating the science of wearable foundation models
arXiv cs.LG
source ↗
Jul 21
Accurate and efficient long-term memory for LLM agents
arXiv cs.AI
source ↗
Jul 21
Multi-level context modeling for consistent expert selection in Mixture-of-Experts
arXiv cs.CL
source ↗
Jul 21
Committed before reasoning: evidence of answer pre-commitment in an open-weight LLM with behavioral reproduction and preliminary activation-level findings
arXiv cs.CL
source ↗
Jul 21
Generalist AI control: a learning-based controller handles systems with varying orders and dynamics using a dynamic state-space representation with attention and masking
arXiv cs.AI
source ↗
Jul 21
BLAD: a multilingual dataset of 1,484 Bangladeshi legal acts from 1799 to 2025
arXiv cs.CL
source ↗
Jul 21
Reinforcement Learning-guided NSGA-II with gray relational coefficient for multi-objective optimization: application to NASDAQ portfolio optimization
arXiv cs.LG
source ↗
Jul 21
PlanFlip: attackers inject prompts in the planning phase to corrupt multi-agent LLM sub-tasks via planning-phase prompt injection
arXiv cs.AI
source ↗
Jul 21
OpenLanguageModel: readable, composable small-language-model pretraining for education and research
arXiv cs.CL
source ↗
Jul 21
Interactive task alignment framed as a POMDP to handle ambiguous user goals in task execution
arXiv cs.AI
source ↗
Jul 21
Berkeley and Heiserman as an unexhausted architecture for embodied machine intelligence
arXiv cs.AI
source ↗
Jul 21
Generative Ontology Induction: domain-agnostic schema discovery from document corpora using large language models
arXiv cs.AI
source ↗
Jul 21
DocOCR-Eval: a correction-based framework for OCR tool selection without ground truth
arXiv cs.LG
source ↗
Jul 21
Diffusion-corrected autoregressive Fourier neural operator for droplet evolution prediction
arXiv cs.LG
source ↗
Jul 21
HantaWatch: federated learning for hantavirus genomic surveillance
arXiv cs.LG
source ↗
Jul 21
SelKV: selective KV cache merging with per-token merge-or-drop and attention compensation
arXiv cs.AI
source ↗
Jul 21
Conformal prediction for self-correcting scientific generation provides statistical guarantees for scientific reasoning validity through progressive absolute-coherent-factuality validation
arXiv cs.CL
source ↗
Jul 21
Real-world evaluation of an AI agent drafting translational impact summaries
arXiv cs.CL
source ↗
Jul 21
ColGraphRAG: late-interaction evidence retrieval for multimodal GraphRAG
arXiv cs.AI
source ↗
Jul 21
Quantizing recursive reasoning models introduces cumulative bias with per-tensor 4-bit quantization, per the arXiv paper on recursive reasoning blocks
arXiv cs.LG
source ↗
Jul 21
Encoding EEG signals to examine human-like next-word prediction behavior in language models
arXiv cs.CL
source ↗
Jul 21
Token-Level cross-modal transformer with contrastive multi-task learning for breast cancer subtype classification and survival prediction
arXiv cs.LG
source ↗
Jul 21
SpecLA: Efficient speculative decoding for linear-attention models
arXiv cs.CL
source ↗
Jul 21
TRACE: trajectory-based safety patch learning for LLM post-training realignment
arXiv cs.LG
source ↗
Jul 21
Some large language models exhibit consistent risk attitudes
arXiv cs.AI
source ↗
Jul 21
AI_LectureNote: a retrospective pilot study of a post-ASR workflow for English-script rendering and semantic drift in Korean-English medical lectures
arXiv cs.CL
source ↗
Jul 21
The failures of marginal influence-based attribution methods for global time series explanations
arXiv cs.LG
source ↗
Jul 20
Loopie: a pair of Mixture-of-Experts looped transformers claim “most powerful looped Transformer to date” with 20B and 6B parameter variants
arXiv cs.CL
source ↗
Jul 20
DSWorld: a data science world model anticipates effects of data science operations to enable efficient autonomous agents
arXiv cs.AI
source ↗
Jul 20
From black box to executable logic: explainable reinforcement learning via Prolog expert systems
arXiv cs.AI
source ↗
Jul 20
Looped Latent Attention compresses K and V caches in looped transformers through a post-training codec, enabling compact cross-loop data storage
arXiv cs.LG
source ↗
Jul 20
ToolVerse enables large-scale environments and long-horizon tasks for agentic reinforcement learning
arXiv cs.AI
source ↗
Jul 20
Harmonizing AI safety thresholds
arXiv cs.AI
source ↗
Jul 20
CAMMAR: Culture-Aware Matryoshka for Metaphorical Arabic Representations
arXiv cs.CL
source ↗
Jul 20
S1-Omni is a unified multimodal reasoning model for scientific understanding, prediction, and generation.
arXiv cs.AI
source ↗
Jul 20
AgentFAIR: a multi-agent framework for FAIRness evaluation of geospatial datasets
arXiv cs.AI
source ↗
Jul 20
Independent certification for trustworthy AI could close the trust gap, argues for a formal certification framework
arXiv cs.AI
source ↗
Jul 20
Adaptive multi-step lookahead decoding for diffusion language models
arXiv cs.CL
source ↗
Jul 20
AnovaX: a local, multi-agent voice assistant with LLM planning, typed executors, and adaptive recovery
arXiv cs.AI
source ↗
Jul 20
Plan, Learn, Adapt (PLA) framework for personalized on-device itinerary generation
arXiv cs.LG
source ↗
Jul 20
Inpainting insights: elevating visual XAI with photorealistic perturbations
arXiv cs.LG
source ↗
Jul 20
Candidate Attended Dialogue State Tracking using BERT
arXiv cs.CL
source ↗
Jul 20
Deep Learning approaches for sleep apnea classification from polysomnographic EEG signals
arXiv cs.LG
source ↗
Jul 20
Causal-Audit: Explicit and auditable graph-based reasoning via target-aware causal chain construction
arXiv cs.AI
source ↗
Jul 20
SkillCorpus aims to consolidate and evaluate the open skill ecosystem for real-world LLM agents
arXiv cs.CL
source ↗
Jul 20
qZACH-ViT: Quantization-Aware Intrinsic Explanations with Recursive Attribution-Stabilized Optimization
arXiv cs.LG
source ↗
Jul 20
Information-Directed Sampling for causal bandits
arXiv cs.LG
source ↗
Jul 20
Toxicity priors show conditional reliability in multilingual and code-mixed abuse detection; English toxicity, Indic abuse, and rule-based severity cues help only in certain linguistic and abuse contexts (ToxGate).
arXiv cs.CL
source ↗
Jul 20
Large language models as unified multimodal learners for clinical prediction
arXiv cs.CL
source ↗
Jul 20
Before the action: benchmarking LLMs on prospective hypothesis discovery
arXiv cs.CL
source ↗
Jul 20
Regularity-aware stochastic MGDA with adaptive conflict-avoidant update direction control
arXiv cs.LG
source ↗
Jul 20
EpiNarrate: agentic generation of grounded narratives from epidemiological scenario projections
arXiv cs.CL
source ↗
Jul 20
From hyperplanes to hyperellipsoids: comparing interpretability of linear and single-qubit mixed-state binary classifiers
arXiv cs.LG
source ↗
Jul 20
Who became financially vulnerable after COVID-19? A population-level machine learning analysis using MEPS data
arXiv cs.LG
source ↗
Jul 20
Structure of the circular-dyadic convolution error
arXiv cs.LG
source ↗
Jul 20
Cura 1T: a specialized model for agentic healthcare
arXiv cs.AI
source ↗
Jul 20
Do coding agents need executable world models, simplification, and verification to solve ARC-AGI-3?
arXiv cs.AI
source ↗
Jul 20
DECODEM: Data extraction from corporate organizational documents via enhanced methods
arXiv cs.CL
source ↗
Jul 20
Frontier AI performance across business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
arXiv cs.CL
source ↗
Jul 20
On the structure of address in multi-party dialogue: moving from discrete labels to continuous levels
arXiv cs.CL
source ↗
Jul 20
Contextual semantic relevance tracks fMRI BOLD responses during naturalistic speech comprehension
arXiv cs.CL
source ↗
Jul 20
NeurOWL: an LLM-based neural-symbolic framework for incomplete OWL ontology reasoning
arXiv cs.AI
source ↗
Jul 20
DrawingVQA: a real-world benchmark for multimodal reasoning on construction drawings
arXiv cs.AI
source ↗
Jul 20
SciForge: a multimodal AI-native workbench for scientific discovery
arXiv cs.AI
source ↗
Jul 20
GraphDx: a cost-aware knowledge-enhanced multi-agent framework for sequential diagnosis
arXiv cs.AI
source ↗
Jul 20
Critical analysis of trustworthy AI tools, Mark frameworks, and the implementation chasms
arXiv cs.AI
source ↗
Jul 20
MLIR-based compilation method for large language models; AI Trading evaluates LLMs for technical market analysis
arXiv cs.CL
source ↗
Jul 20
Verbalizable representations form a global workspace in language models
arXiv cs.CL
source ↗
Jul 20
SeerGuard: a consequence-aware safety framework for mobile GUI agents using world model prediction
arXiv cs.AI
source ↗
Jul 20
Process reward informed tree rollout for effective multi-turn RL
arXiv cs.CL
source ↗
Jul 20
A transportable threshold-based framework for interpretable classification of medical data
arXiv cs.LG
source ↗
Jul 20
VarRate: Training-free variable-rate KV cache compression for long-context LLMs
arXiv cs.CL
source ↗
Jul 20
Controlling implicit shortcut reliance in L2 spoken English auto-markers
arXiv cs.CL
source ↗
Jul 20
Neuro-symbolic AI for LEED compliance: document-centric benchmarking, deterministic numeric checking, and when multimodal hurts
arXiv cs.AI
source ↗
Jul 20
ADS-C: Antidistillation Sampling for Classification
arXiv cs.LG
source ↗
Jul 20
Quantum program generation should prioritize validity over probabilistic scaling, arguing that scaling parameters to boost emergent reasoning is a directional error for quantum circuit synthesis due to strict mathematical constraints and a syntax-semantics gap.
arXiv cs.LG
source ↗
Jul 20
Robust peak-cost constrained reinforcement learning
arXiv cs.LG
source ↗
Jul 20
Stochastic Reset Pathfinding: path-level regret for cascading bandits over graph paths
arXiv cs.LG
source ↗
Jul 20
Behavioral controllability of agentic models for information extraction: from fixed workflows to reflective agents
arXiv cs.AI
source ↗
Jul 20
From plausible to actionable: a position on LLM self-explanations
arXiv cs.CL
source ↗
Jul 20
How much human label variation does formal semantic structure explain? Group-level effects and item-level ceilings in NLI
arXiv cs.CL
source ↗
Jul 20
Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
arXiv cs.CL
source ↗
Jul 20
Precise but uncoupled: reviewer precision does not guarantee critique uptake in multi-agent math reasoning
arXiv cs.AI
source ↗
Jul 20
A Formally Grounded ODRL Evaluator: Implementation and Comparison
arXiv cs.AI
source ↗
Jul 20
BayesPO: Bayesian prompt optimization via parallel-tempered gradient-guided discrete MCMC
arXiv cs.CL
source ↗
Jul 20
Publicly-Verifiable Certificates for Statistical Algorithms show non-interactive proofs of learning with public, distributionally robust certification of a hypothesis's validity.
arXiv cs.LG
source ↗
Jul 20
Kolmogorov–Arnold Networks replace fixed activations with learned edge functions in a six-layer, 10M-parameter B-spline model; study reports high reconstruction of feed-forward edges and effects of pruning.
arXiv cs.LG
source ↗
Jul 20
Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention
arXiv cs.LG
source ↗
Jul 20
Better starts, better ends: bootstrapped iterative self-reasoning distillation for compressed reasoning
arXiv cs.CL
source ↗
Jul 20
Beyond a joke: multi-angle reasoning for detecting and explaining harmful humor in memes
arXiv cs.AI
source ↗
Jul 20
Harness-in-the-loop learning: using data-generating harnesses to shape future foundation models through execution traces
arXiv cs.LG
source ↗
Jul 20
Auto-scaling approach for serverless environments using a multi-expert consensus mechanism
arXiv cs.LG
source ↗
Jul 20
Frontier language models struggle to copy text; 2D viewing improves assessment of copying behavior
arXiv cs.CL
source ↗
Jul 20
Diffusion models recover accurate mixture weights despite score function insensitivity
arXiv cs.LG
source ↗
Jul 20
MGDT: MLLM-guided diffusion transformer with relation-adaptive mixture-of-experts for multimodal knowledge graph completion
arXiv cs.AI
source ↗
Jul 20
Logic, optimization, and artificial intelligence; logic and optimization combine to enhance rule-based AI for transparency and reproducibility
arXiv cs.AI
source ↗
Jul 20
Knowledge-centric agents for workflow generation
arXiv cs.AI
source ↗
Jul 17
MIT probe prompts questions about belonging and problem-solving; assistant professor Bailey Flanigan reflects on a sense of belonging at MIT
MIT News (AI)
source ↗
Jul 17
Explainability research should prioritize foundations over ad-hoc methods; position paper argues for integrating explanations into end-to-end, human-in-the-loop systems.
arXiv cs.LG
source ↗
Jul 17
Value leakage: a language model's answers are shaped by its own values
arXiv cs.LG
source ↗
Jul 17
Just Keep Prompting: evaluating repetitive Socratic prompting in VLMs
arXiv cs.CL
source ↗
Jul 17
Quantum compositional NLP for Arabic applies pregroup grammar to map sentences to quantum circuits and mirror grammatical structure; Arabic morphosyntax tested in quantum topology
arXiv cs.CL
source ↗
Jul 17
LBA proposes textual hard-label adversarial attack under low query budgets
arXiv cs.CL
source ↗
Jul 17
UniSAGE unifies static and dynamic attributes with hyper-structure for hierarchical data modeling
arXiv cs.CL
source ↗
Jul 17
Latent communication between language model agents: channels, alignment, and the limits of text
arXiv cs.CL
source ↗
Jul 17
UzWordnet and Generative AI enable Uzbek language learning through game-playing
arXiv cs.CL
source ↗
Jul 17
Automatically evolving prompt guidelines for task-specific optimization
arXiv cs.CL
source ↗
Jul 17
Token Time Continuous Diffusion for Language Modeling introduces a diffusion language model that maps Gaussian noise to a token canvas in continuous space with per-token time steps.
arXiv cs.CL
source ↗
Jul 17
Eta Given Delta: Defining LLM tool efficiency with marginal tool utility
arXiv cs.CL
source ↗
Jul 17
Simplicity paradox: debunking myths about prompting and datasets for LLM evaluation
arXiv cs.CL
source ↗
Jul 17
MAPS models co-existing subjective perspectives and shared meaning in multi-agent cognitive dialogue
arXiv cs.CL
source ↗
Jul 17
Introspection fine-tuning (IFT): training small LLMs to introspect
arXiv cs.CL
source ↗
Jul 17
arXiv:2607.14112v1 paper shows language models have reliability ceilings from resolvable output uncertainty limits
arXiv cs.CL
source ↗
Jul 17
T5-CSBoost: adversarial perturbation resistant LLM fingerprinting
arXiv cs.CL
source ↗
Jul 17
CoEvoT introduces co-evolving chain-of-thought prompting for graph-LLM reasoning under distribution shift.
arXiv cs.CL
source ↗
Jul 17
ReportMedSAM: guiding segmentation through radiology reports
arXiv cs.CL
source ↗
Jul 17
Heterogeneous element-aware cross-version differencing of scientific documents via layout-aware alignment and structure-aware reasoning
arXiv cs.CL
source ↗
Jul 17
Budgeted subset refinement for execution-aware LLM research ideation
arXiv cs.CL
source ↗
Jul 17
Semantic register compression in multi-agent LLM cascades.
arXiv cs.CL
source ↗
Jul 17
Study presents first cross-dataset Urdu fake news detection using XLM-RoBERTa; length confound analysis included.
arXiv cs.CL
source ↗
Jul 17
refusal in the first half: a mechanistic study of the prefill jailbreak; four models show harm remains represented and refusal behavior drops to chance after a prefill prompt
arXiv cs.CL
source ↗
Jul 17
Study introduces Implicit Reasoning Steering via Concept Chaining to address LLM reasoning fragility
arXiv cs.CL
source ↗
Jul 17
Severance problem: LLMs are unaware of the person beyond the prompt
arXiv cs.CL
source ↗
Jul 17
CARPRT: Class-aware zero-shot prompt reweighting for black-box vision-language models.
arXiv cs.LG
source ↗
Jul 17
Explainable geospatial AI for satellite ground station siting using LiDAR-derived terrain intelligence
arXiv cs.LG
source ↗
Jul 17
Certified domain consistency for multi-domain retrieval with label-free per-domain contamination control and conformal risk guarantees
arXiv cs.LG
source ↗
Jul 17
QFireNet: Quantum-Enhanced U-Net for Wildfire Segmentation from Sentinel-2 Imagery
arXiv cs.LG
source ↗
Jul 17
Branching policy optimization for sandbox-native language agent reinforcement learning
arXiv cs.LG
source ↗
Jul 17
arXiv:2607.14174v1 study extends supervised lexicon-learning to 10-K filings and Item 1A risk-factor sections; trains sentiment scores on return and volatility labels at multiple aggregation levels
arXiv cs.LG
source ↗
Jul 17
Low-latency relay selection in NR-V2X vehicular networks using graph isomorphism networks with edge features
arXiv cs.LG
source ↗
Jul 17
Renew: Towards learning world models and repairing model exploitation from preferences
arXiv cs.LG
source ↗
Jul 17
Closed-loop knowledge dynamics saturate under internal feedback; external information can move knowledge states beyond current attractors, using a three-level operational framework with transition kernels.
arXiv cs.LG
source ↗
Jul 17
Time-to-event model using a temporal machine learning framework to predict ALS progression and healthcare utilization
arXiv cs.LG
source ↗
Jul 17
TEDDY, a decoder transformer trained on ICD-10 diagnoses from 1.6 million children's EHRs, models pediatric health risks.
arXiv cs.LG
source ↗
Jul 17
Long-term user engagement optimization through model-agnostic downstream rewards learning
arXiv cs.LG
source ↗
Jul 17
Researchers propose augmentations for robust and efficient imitation learning in streamed video games.
arXiv cs.LG
source ↗
Jul 17
Study finds privacy risks in federated learning for radiology reports; analyzing tokenizer-driven leakage via gradient inversion
arXiv cs.LG
source ↗
Jul 17
LIGO-PINN: learned initialization via gated optimization to alleviate convergence failures in physics informed neural networks
arXiv cs.LG
source ↗
Jul 17
MIDiff: Tackling sparsity and imbalance in mobile usage generation via multivariate-imaging diffusion
arXiv cs.LG
source ↗
Jul 17
Local additive feature attribution: a mathematical taxonomy and reporting checklist
arXiv cs.LG
source ↗
Jul 17
Lyapunov-guided framework LyaGuide enables stabilization of generative flows with post-training guidance
arXiv cs.LG
source ↗
Jul 17
NeuroGRIP: retrieval-augmented graph refinement for knowledge-grounded EEG seizure diagnosis
arXiv cs.LG
source ↗
Jul 17
COAT framework learns interpretable prescriptive policies from observational data using counterfactual estimates and mixed-integer optimization for constrained decision making
arXiv cs.LG
source ↗
Jul 17
Policy learning with missing treatment data; effects on average treatment effect and policy value estimation.
arXiv cs.LG
source ↗
Jul 17
Dysco: Dynamic Subspace Boosting to mitigate LoRA interference in federated learning
arXiv cs.LG
source ↗
Jul 16
Graded entity-familiarity readouts in language models: Polish adaptation, cross-language robustness, and refusal steering
arXiv cs.CL
source ↗
Jul 16
Text2Sign, a text-to-sign language video diffusion model, runs on a single NVIDIA L4 GPU as a baseline.
arXiv cs.CL
source ↗
Jul 16
Explaining reinforcement learning agents via inductive logic programming
arXiv cs.AI
source ↗
Jul 16
New paper explores federated learning and explainable AI integration for privacy-preserving transparent machine learning
arXiv cs.LG
source ↗
Jul 16
OriginBlame introduces record- and token-level data provenance for AI training datasets
arXiv cs.AI
source ↗
Jul 16
SPINE: bridging the cyber-physical gap with agentic AI
arXiv cs.AI
source ↗
Jul 16
Researchers introduce black-box interventional grounding audits to test LLM premise dependency via predicate substitution; method replaces target predicates with fresh symbols and re-runs models to assess reasoning changes.
arXiv cs.AI
source ↗
Jul 16
Probabilistic extension of neuro-symbolic AGI robots based on Belnap's typed intensional FOL
arXiv cs.AI
source ↗
Jul 16
Survey Examines Self-Improving Agentic Systems' Shift to Deployment
arXiv cs.AI
source ↗
Jul 16
Improving molecular property prediction in small language models using graph-based tools.
arXiv cs.AI
source ↗
Jul 16
Oracle Agent Memory as an enterprise memory substrate for long-horizon AI agents
arXiv cs.AI
source ↗
Jul 16
Researchers introduce DROPJ, a human-centered method for safe agent training and deployment in safety-critical environments with unknown dynamics and no reward function.
arXiv cs.AI
source ↗
Jul 16
CayleyR solves TopSpin puzzle by detecting cycle intersections in Cayley graphs
arXiv cs.AI
source ↗
Jul 16
Networked intelligence: active shared context graphs for human-AI team science
arXiv cs.AI
source ↗
Jul 16
AI-native insurance for agentic AI: pricing, underwriting, and end-to-end automation
arXiv cs.AI
source ↗
Jul 16
Cost-optimal foundation model deployment portfolio for transportation management
arXiv cs.AI
source ↗
Jul 16
Harness Handbook: making evolving agent harnesses readable, navigable, and editable
arXiv cs.AI
source ↗
Jul 16
Theory-Level Autoformalization: from isolated statements to unified formal knowledge bases
arXiv cs.AI
source ↗
Jul 16
Researchers present EZSMT Version 3, an extensible SMT-based CASP framework.
arXiv cs.AI
source ↗
Jul 16
Set-shifting benchmark tests how LLM agents adapt to hidden reliability shifts in tool availability
arXiv cs.AI
source ↗
Jul 16
LAPO uses leave-one-turn attribution for multi-turn search reasoning as self-generated process supervision.
arXiv cs.AI
source ↗
Jul 16
OpenRCA dataset highlights challenges in achieving high accuracy for root cause analysis on real-world telemetry data.
arXiv cs.AI
source ↗
Jul 16
Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling
arXiv cs.AI
source ↗
Jul 16
AI advice reduces willingness to say “I don’t know,” even when the AI is wrong and accuracy is incentivized
arXiv cs.AI
source ↗
Jul 16
Safety Sentry: Context-aware human intervention via execute-ask-refuse routing
arXiv cs.AI
source ↗
Jul 16
LLM-powered agentic system enables automatic discovery of ordinary differential equations in biological systems
arXiv cs.AI
source ↗
Jul 16
Stocktake: Measuring the gap between perception and action in LLM agents with a fair oracle
arXiv cs.AI
source ↗
Jul 16
UESF-Bench benchmarks unified embodied seeking and following, addressing existing benchmarks' assumption of initial target visibility.
arXiv cs.AI
source ↗
Jul 16
FixItFlow automates troubleshooting guide generation from historical cloud incidents using large language models
arXiv cs.CL
source ↗
Jul 16
Safe-Psych, a sequential evaluation benchmark for LLMs in psychiatry, addresses incomplete information by requiring clarification.
arXiv cs.CL
source ↗
Jul 16
Do LLMs need architectural changes for simultaneous speech translation? A prefix-to-prefix data driven approach
arXiv cs.CL
source ↗
Jul 16
Researchers use persona vectors to audit open-weight LLMs, revealing 53 traits across four models in arXiv:2607.13162v1.
arXiv cs.CL
source ↗
Jul 16
RAGthoven presents multi-stage LLM pipeline for SemEval-2026 Task 1's multilingual humor generation in English, Spanish, and Chinese.
arXiv cs.CL
source ↗
Jul 16
Adaptive Filtering of the KV cache: diagnosing and correcting structural-role bias in LLM inference
arXiv cs.CL
source ↗
Jul 16
Gsm-Plus-Bn Introduces Perturbation-Based Benchmark for Bangla Mathematical Reasoning in LLMs; Addresses Lack of Evaluation in Bengali
arXiv cs.CL
source ↗
Jul 16
Discourse-aware policy analysis with argumentation: a hybrid LLM-symbolic framework for disaster governance
arXiv cs.CL
source ↗
Jul 16
Finding the right tables and columns: a benchmark and corpus-adaptive embeddings for SQL schema retrieval
arXiv cs.CL
source ↗
Jul 16
Researchers propose a meta-learning framework to address data scarcity in low-resource language alignment for multilingual LLMs; arXiv:2607.13315v1
arXiv cs.CL
source ↗
Jul 16
Evaluation ability does not imply optimization utility; LLM-as-a-judge signals in closed-loop table recognition show weak judge signals on FinTabNet and OmniDocBench
arXiv cs.CL
source ↗
Jul 16
A POS tier automates annotation for low-resource language documentation; neural interlinear glossing for Irabu (Southern Ryukyuan) presented
arXiv cs.CL
source ↗
Jul 16
GFlowRL: scaling distribution-matching RL to large language models
arXiv cs.CL
source ↗
Jul 16
Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
arXiv cs.CL
source ↗
Jul 16
Small language models for biomedical data-to-text generation: a case study on medication leaflets and post-training alignment methods
arXiv cs.CL
source ↗
Jul 16
Cross-rubric generalization: training on one rubric set and evaluating on a different rubric set for critical thinking essay scoring
arXiv cs.CL
source ↗
Jul 16
Researchers introduce a benchmark and reference system for live captioning Sikh Kirtan; arXiv:2607.13457v1 presents a closed-vocabulary task requiring exact SGGS line transcription.
arXiv cs.CL
source ↗
Jul 16
Benchmarking cross-device agents in heterogeneous environments
arXiv cs.CL
source ↗
Jul 16
MyAG is a graph-based framework for designing and analyzing composable LLM agent systems with three graph abstractions.
arXiv cs.CL
source ↗
Jul 16
Study introduces BioASQ Task 14B 2026 system using hybrid retrieval and multi-model answer combination.
arXiv cs.CL
source ↗
Jul 16
Memory as a controlled process: learned adaptive memory management for LLM agents
arXiv cs.CL
source ↗
Jul 16
Analogical Deep Research: retrieving and integrating historical analogies for foresight analysis
arXiv cs.CL
source ↗
Jul 16
Automatic differentiation from scratch: how PyTorch computes gradients for physics-informed neural networks
arXiv cs.LG
source ↗
Jul 16
Researchers propose lightweight training strategy for efficient transfer learning by decoupling feature extraction and classifier optimization with margin-based weighted loss.
arXiv cs.LG
source ↗
Jul 16
Masking, fingerprinting, and privacy from discarded geometry: what your model threw away and why you'll want it back
arXiv cs.LG
source ↗
Jul 16
Targeted PD identifies only the components that process specific inputs to scale parameter decomposition in neural networks.
arXiv cs.LG
source ↗
Jul 16
Uncertainty-aware sequential decision rules for event-triggered LLM invocation in streaming systems
arXiv cs.LG
source ↗
Jul 16
TSSM model enhances global station weather forecasting by addressing accuracy limitations through temporal-variable-historical modeling.
arXiv cs.LG
source ↗
Jul 16
Disentangling knowledge states with ability and proficiency modeling for knowledge tracing
arXiv cs.LG
source ↗
Jul 16
STKAN: Kolmogorov-Arnold networks for spatio-temporal forecasting
arXiv cs.LG
source ↗
Jul 16
Researchers propose Samba, a hybrid Mamba model for audio-visual navigation, addressing inadequacies of existing frameworks since 2020.
arXiv cs.LG
source ↗
Jul 16
Researchers introduce CoDiffGRN, a method for gene regulatory network inference using BEELINE-KGC benchmark and co-evolutionary discrete diffusion.
arXiv cs.LG
source ↗
Jul 16
ShortOPD: recovering pruned LLMs with short-to-long on-policy distillation
arXiv cs.LG
source ↗
Jul 16
Hedgehog: hierarchical evaluation of drug generators through rigorous filtration
arXiv cs.LG
source ↗
Jul 16
SteinGate: tail-sensitive safe reinforcement learning via Stein discrepancy
arXiv cs.LG
source ↗
Jul 16
Concurrent image understanding and generation via self-correcting coupled Markov jump processes
arXiv cs.LG
source ↗
Jul 16
EMAGN uses learned clustering for scalable traffic forecasting.
arXiv cs.LG
source ↗
Jul 16
Muon emerges as strong optimizer for deep learning, outperforming Adam and AdamW; theoretical work interprets it as steepest descent under spectral norm.
arXiv cs.LG
source ↗
Jul 16
Deconstructing Actor-Critic: a large-scale empirical study of design components for practitioners
arXiv cs.LG
source ↗
Jul 16
Tabular foundation models for discrete choice estimation show limited performance due to row-independence assumptions
arXiv cs.LG
source ↗
Jul 16
Accuracy-preserving stability regularization for large-scale retail demand forecasting
arXiv cs.LG
source ↗
Jul 16
Agora: collective and permissionless internet-scale pretraining of large language models
arXiv cs.LG
source ↗
Jul 16
Weight feedback computes the Jacobian transpose locally in modern deep networks
arXiv cs.LG
source ↗
Jul 16
Where should RL post-training compute go? Model size, search, learning, and feedback
arXiv cs.LG
source ↗
Jul 16
arXiv:2607.13395v1: Enlightenment-style post-tuning without training for large models inspired by human 'enlightenment' concept
arXiv cs.LG
source ↗
Jul 16
Study compares KANs and MLPs on structured data classification using twelve datasets.
arXiv cs.LG
source ↗
Jul 16
GIFT teaches vision-language generative AI models to produce CAD programs for simulating and testing 3D objects.
MIT News (AI)
source ↗
Jul 15
3 questions: neural transparency and the future of AI design
MIT News (AI)
source ↗
Jul 15
Towards demystifying the creativity of diffusion models
Google Research
source ↗
Jul 14
Devavrat Shah develops system that uses tabular data for real-time planning at scale
MIT News (AI)
source ↗
Jul 14
JARVIS Challenge tests AI copilots in tough-tech engineering for jet engine design
MIT News (AI)
source ↗
Jul 13
MIT students help prevent cyberattacks at Bridgewater State University cyber range and security operations center
MIT News (AI)
source ↗
Jul 13
MIT researchers use three AI agents to create 3D scene recreations for robot training data
MIT News (AI)
source ↗
Jul 13
Researchers develop an evaluation procedure to test generative AI models for harmful capabilities without generating outputs, aiding auditors in identifying open-source models adapted to produce illegal content.
MIT News (AI)
source ↗
Jul 9
Aurora 1.5 adds 22 variables, hourly resolution, and probabilistic ensembles to extend open foundation models for weather and Earth-system applications.
Microsoft Research
source ↗
Jul 9
Tiny robot boats assemble floating structures and reconfigure themselves with minimal human direction
MIT News (AI)
source ↗
Jul 9
SensorFM aims for a general intelligence and interface for wearable health data
Google Research
source ↗
Jul 8
Flint: a visualization language for the AI era
Microsoft Research
source ↗
Jul 7
How novice coders can develop AI programs for military applications
MIT News (AI)
source ↗
Jul 7
Jesse Thaler named director of the Laboratory for Nuclear Science
MIT News (AI)
source ↗
Jul 7
AI model inference costs fall sharply; GPT-4-class costs drop from ~$30 per million tokens in 2023 to under $1 today
BAIR (Berkeley)
source ↗
Jul 6
Envisioning the future of computing prize submission titled "Superintelligence, Superintimate" by Rachel Sava; aims to preserve benefits of neurotechnology for all
MIT News (AI)
source ↗
Jul 1
MIT in the media: President Sally Kornbluth and ASU President Michael Crow discuss training the next generation of scientists for America’s evolving tech landscape
MIT News (AI)
source ↗
Jul 1
BAIR Graduate Showcase recognizes Ph.D. graduates from the Berkeley Artificial Intelligence Research Lab, class of 2026.
BAIR (Berkeley)
source ↗
← 2026-08
2026-06 →