LIVE · refreshes every 20 min
updated Sep 16, 03:23 UTC
110010
.art
(ificial intelligence)
AI products, research, and launches — clustered and ranked, not just re-blogged.
All stories
Research
Discussion
RSS
⌕
Search
/
☾
110010
.art
/ archive
Research
Aug 27
19d ago
FrontierChallenge evaluates 300 end-to-end scientific workflows; authors release and evaluate 97 tasks across multiple domains
arXiv cs.AI
→ story
19d ago
Function-level execution feedback for code preference optimization
arXiv cs.AI
→ story
19d ago
FuzzingBrain-Bench V1 evaluates open-ended bug discovery by LLMs
arXiv cs.AI
→ story
19d ago
Gated activation steering reduces sycophancy and hallucination in medical question answering
arXiv cs.AI
→ story
19d ago
Generating biomedical fact-checking reports with RL-enhanced agentic search
arXiv cs.AI
→ story
19d ago
Granite.Trust policy tools: shareable, actionable policies for generative AI applications
arXiv cs.AI
→ story
19d ago
GRAPE: Gradient refinement and progress-aware exploitation for query-efficient high-dimensional Bayesian optimization
arXiv cs.LG
→ story
19d ago
GreenLeaf Law Embed Tiny achieves competitive performance with a 0.6B parameter embedding model for legal domain retrieval
arXiv cs.LG
→ story
19d ago
Groundhog bit-flip attack: seeding infinite generation loops in mixture-of-experts LLMs through bit flips
arXiv cs.CL
→ story
19d ago
GUIDE: Generative unsupervised Chinese query correction via phonetic and visual shared-ID encoding
arXiv cs.CL
→ story
19d ago
HealthBench-Psych: a mental health subset of OpenAI's HealthBench
arXiv cs.CL
→ story
19d ago
Hierarchical MoE for multi-modal ILD diagnosis
arXiv cs.AI
→ story
19d ago
How much of a measured AI preference is the model and how much is the instrument; multiple instruments disagree on the source of model preferences
arXiv cs.AI
→ story
19d ago
In-Context Inpainting for Time Series Forecasting reframes forecasting as visual inpainting using large vision models.
arXiv cs.AI
→ story
19d ago
LifePlanner benchmarks LLM agents for geo-spatial planning using map data enriched with large-scale local social media posts.
arXiv cs.AI
→ story
19d ago
LLM agents perform controlled experiments using simulation models
arXiv cs.AI
→ story
19d ago
LLM-driven, datasheet-aware framework for early-stage hardware compatibility verification identifies documentation-level interface incompatibilities from datasheets and connectivity descriptions
arXiv cs.AI
→ story
19d ago
LLMs trained on synthetic limit order book data generate valid LOB event sequences but fail to learn the LOB state, causing biased estimates and spurious predictability in forecasting future LOB events.
arXiv cs.AI
→ story
19d ago
MacroAgent: Regularity-aware macro legalization with LLM-agent-designed contour algorithms
arXiv cs.LG
→ story
19d ago
Measurement-budget allocation in quantum learning with finite-shot generalization guarantees
arXiv cs.AI
→ story
19d ago
Modeling pragmatic representations of conversations improves cross-domain derailment forecasting in low-data settings
arXiv cs.CL
→ story
19d ago
MSR-IVA: Masked Structural Residual Independent Vector Analysis for state-aware fusion of structural MRI and dynamic functional network connectivity
arXiv cs.LG
→ story
19d ago
MTDiag: a multi-turn diagnostic dataset for clinically meaningful LLM evaluation
arXiv cs.CL
→ story
19d ago
Multi-modal anomaly detection: a survey
arXiv cs.LG
→ story
19d ago
Natural Language Input, semantic track representation, and LLM inference: making the Maritime Information Exchange Model tractable
arXiv cs.AI
→ story
19d ago
NVExplain: Explaining time series forecasting with latent trajectory analysis and structure-preserving surrogates
arXiv cs.LG
→ story
19d ago
Padamitra: grounded glossary generation for classical Sanskrit
arXiv cs.CL
→ story
19d ago
PhaseShift harmonizes heterogeneous traffic trajectories into a shared representation and trains a reusable backbone for topology-aware data consolidation across signalized intersections
arXiv cs.AI
→ story
19d ago
PhysElite: evaluating how far LLMs are from solving olympiad-level physics problems
arXiv cs.AI
→ story
19d ago
PostgreSQL-native graph RAG engine for retrieval-augmented generation using graph-based retrieval
arXiv cs.AI
→ story
19d ago
Reliable LLM-powered decision engines for large-scale supply chain operations: architecture, safety, and performance guarantees
arXiv cs.AI
→ story
19d ago
Relieving RAG bottlenecks via evidence frontloading and pressure-adaptive budgeting
arXiv cs.CL
→ story
19d ago
RENDER: Controlling reader-facing evidence in LLM memory evaluation
arXiv cs.AI
→ story
19d ago
Representational geometry of dynamic programs described as shortest paths on DAGs, tropical polynomials, and Newton polyhedra isomorphic as semirings
arXiv cs.LG
→ story
19d ago
Resource-Efficient pruning for transformer via low-rank importance estimation
arXiv cs.LG
→ story
19d ago
Retrieve, match, escalate: accurate and scalable product linking with VLM-distilled cross-encoders and agentic VLMs
arXiv cs.AI
→ story
19d ago
Rollout-Decoded Reconstruction for long-horizon prediction in latent world models
arXiv cs.LG
→ story
19d ago
Routed Graph Handoff: adaptive format selection for multi-agent LLM delegation
arXiv cs.CL
→ story
19d ago
SelfGraphRAG: bridge the supervision gap in graph-based RAG with synthetic QA generation
arXiv cs.CL
→ story
19d ago
Semantic variability of replies across LLMs; implications for designing conversation-based assessment
arXiv cs.CL
→ story
←
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
→
Older items: monthly archive →