本篇博文主要内容为 2026-07-28 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。

说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。

提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。

目录

概览 (2026-07-28)

今日共更新1242篇论文,其中:

  • 自然语言处理141篇(Computation and Language (cs.CL))
  • 人工智能406篇(Artificial Intelligence (cs.AI))
  • 计算机视觉244篇(Computer Vision and Pattern Recognition (cs.CV))
  • 机器学习326篇(Machine Learning (cs.LG))
  • 多智能体系统21篇(Multiagent Systems (cs.MA))
  • 信息检索47篇(Information Retrieval (cs.IR))
  • 人机交互45篇(Human-Computer Interaction (cs.HC))

多智能体系统

[MA-0] Decentralised Consensus Learning Networks: SME Rotation Without Centralised Reward

【速读】:该论文旨在解决现代人工智能学习系统中依赖中心化奖励信号所带来的局限性,即强制赋予单一外部定义的正确性或价值标准,从而抑制了知识的多样性与自主演化。其核心问题是:如何在缺乏预设权威或地面真值的情况下,实现专家能力的自组织涌现与动态分配。解决方案的关键在于提出一种去中心化、基于共识的多智能体学习框架,其中智能体通过加权社会共识更新信念,信任度依据同伴一致性推断出的胜任力进行分配,而非依赖于真实标签。主题专家(SME)身份以动态的前百分位胜任力排名形式授予,而非固定标签。实验结果表明,该框架在多种网络拓扑、规模及信念表示下均表现出鲁棒性,尤其在高维向量信念空间中,随着维度增加,系统自发形成局部共识并导致专家地位高度集中于单一智能体,这一现象被证实由信念维度本身驱动,而非随机噪声,揭示了去中心化学习在复杂高维共识空间中的涌现特性:最持续与集体信念一致的智能体自然成为被认可的专家。

链接: https://arxiv.org/abs/2607.24416
作者: Florin Neagu
机构: 未知
类目: Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Centralised reward signals dominate modern AI learning systems, but they impose a single external definition of correct or valuable knowledge. We present a decentralised, consensus-based multi-agent learning framework in which expertise emerges through peer validation rather than prescribed reward. Agents update beliefs via weighted social consensus, while trust is allocated according to competence inferred from peer consistency instead of ground truth. Subject-matter expert (SME) status is assigned dynamically as a top-percentile competence rank rather than a fixed label. We evaluate the framework across 84 simulation runs spanning 30 to 10,000 agents, multiple graph topologies, sparse large-scale networks, scalar and vector belief representations, dimensionality sweeps (D=1-500), multi-seed robustness tests, and parameter sensitivity analyses. Phase 1 shows that SME rotation is robust, persistent, topology-invariant, and scale-invariant: 90-100% of agents attain SME status, with most expertise turnover occurring after belief convergence and increasing with network size. Phases 2 and 3 show that vector beliefs introduce heterogeneous convergence with cascade dynamics and reveal five distinct dynamical regimes as belief dimensionality increases. At high dimensionality (D=150-200), the network reaches stable partial consensus while expertise becomes increasingly concentrated in a single agent. ETA sensitivity analysis demonstrates that this concentration is driven by belief dimensionality rather than stochastic noise. We interpret this behaviour as an emergent property of decentralised learning: in complex high-dimensional consensus spaces, the agent most consistently aligned with the collective belief naturally emerges as the recognised expert. Subjects: Multiagent Systems (cs.MA) Cite as: arXiv:2607.24416 [cs.MA] (or arXiv:2607.24416v1 [cs.MA] for this version) https://doi.org/10.48550/arXiv.2607.24416 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[MA-1] Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG

【速读】:该论文旨在解决海洋生物监测中面临的严苛能耗约束、水下通信能力差以及远程部署设备传输原始多模态数据成本高昂的问题。其核心解决方案在于提出一种低功耗的水下监测架构,通过“始终在线的边缘感知”与“按需激活的高性能本地推理”相结合的方式实现能效与智能分析的平衡。该架构采用分层主-卫星设计:超低功耗的MAX78000/MAX78002微控制器持续采集视觉与声学信号,而仅在预定任务、事件触发或研究人员交互时激活NVIDIA Jetson Orin NX进行本地高阶处理。一旦激活,Jetson执行完整的本地多模态处理流程,包括数据摄入、目标提取、基于嵌入的索引、物种识别、检索增强推理及自动化报告生成。系统利用BioCLIP/OpenCLIP嵌入将任务数据、海洋分类参考、科学文献和操作元数据组织于本地ChromaDB数据库中,并通过融合视觉相似性搜索、质心分类与监督分类器的专用识别层实现自适应物种识别。同时,基于LangChain的多智能体框架负责查询路由、结构化分析、能耗管理、硬件重构及报告生成。该架构在视觉与声学监测案例中进行了验证,成功实现了从超低功耗连续感知到本地多模态智能的无缝衔接,使水下观测站能够生成结构化、可研究的知识输出,同时对数据进行压缩以支持灵活的声学、光学或卫星传输,显著降低能源消耗与通信开销。

链接: https://arxiv.org/abs/2607.24313
作者: Mohamed Amine Janati,Laurent Gautier,Stéphane Barbot
机构: 未知
类目: Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV); Multiagent Systems (cs.MA); Robotics (cs.RO)
备注: Contains 25 pages, 11 figures. To be submitted to Elsevier Ocean Engineering

点击查看摘要

Abstract:Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cost of transmitting raw multimodal data from remote deployments. This paper proposes a low-consumption underwater monitoring architecture that combines always-on edge sensing with selective high-performance local reasoning. The system follows a hierarchical master–satellite design in which ultra-low-power MAX78000/MAX78002 microcontrollers continuously monitor visual and acoustic signals, while an NVIDIA Jetson Orin NX is activated only for scheduled processing, event-driven analysis, or researcher interaction. Once active, the Jetson executes a fully local multimodal pipeline for data ingestion, visual target extraction, embedding-based indexing, species identification, retrieval-augmented reasoning, and automated reporting. BioCLIP/OpenCLIP embeddings are used to organize mission data, marine taxonomic references, scientific documents, and operational metadata in local ChromaDB collections. A dedicated identification layer combines visual similarity search, centroid-based classification, and supervised classifiers to support adaptive species recognition. A LangChain-based multi-agent framework coordinates query routing, structured analysis, energy management, hardware reconfiguration, and report generation. The architecture is evaluated through visual and acoustic monitoring case studies. The proposed system bridges ultra-low-power continuous sensing with local multimodal intelligence, enabling underwater stations to produce structured, researcher-ready knowledge while compressing local data for flexible acoustic, optical, or satellite transmission, minimizing both energy use and communication overhead.

[MA-2] Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

【速读】:该论文旨在解决自改进智能体(self-improving agents)在迭代优化过程中因自我验证机制失效而导致的“验证-部署差距”(verifier–deployment gap)问题。具体而言,当智能体同时控制其策略(policy)与验证机制时,其自编写的测试或评分标准可能持续维持高分,但实际部署性能却出现退化或停滞,从而导致不可靠的更新被部署。该问题的核心在于:智能体在反复重写自身策略与验证逻辑的过程中,会逐渐偏离真实部署环境下的目标表现,而这种偏差随智能体能力提升并未消除,反而呈现出能力分层特征——弱智能体倾向于破坏已习得的有效策略,而强智能体虽更稳定,但仍无法准确衡量部署分布。为应对这一挑战,论文提出密封外部接受环(Sealed Exogenous Acceptance Loop, SEAL),其关键在于保留智能体自编测试的同时,引入一个不可篡改、不可观测的外部审计机制(harness-side audit),由第三方固定评估流程对候选策略与当前最优版本进行对比。该审计过程完全独立于智能体控制,仅返回“接受/拒绝”信号,且在检测到明显退化时强制回滚至旧版本状态。实验表明,SEAL在六种模型和三个随机种子下均显著优于无保护基线,证明了即使不放弃自验证机制,只需引入至少一个受控于智能体之外的部署接纳信号,即可实现可靠自改进。

链接: https://arxiv.org/abs/2607.24300
作者: Diandian Guo,Cong Cao,Fangfang Yuan,Yingqi Wang,Yueshan Wang,Dakui Wang
机构: Guangzhou University(广州大学); South China Normal University(华南师范大学)
类目: Computation and Language (cs.CL); Multiagent Systems (cs.MA)
备注: 9 pages, 6 figures

点击查看摘要

Abstract:Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier–deployment gap. This gap refers to the discrepancy between an agent’s self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability, and how little exogenous trust is sufficient to prevent real regressions from being deployed. To address this problem, we introduce a Sealed Exogenous Acceptance Loop (SEAL). SEAL retains self-authored tests but compares each candidate with the incumbent through a fixed harness-side audit. The agent cannot author or inspect the audit, receives only accept/reject, and the whole incumbent state is retained after a clear regression. Our experiments show that this problem often appears in heuristic learning settings. These settings require trial-and-error discovery of the target objective. We further find that failures of self-written verification are stratified by capability. Weaker agents tend to damage previously acquired strategies behind easy self-tests. Stronger agents are more stable, but they still mismeasure the deployment distribution. Standard self-written constraints do not reliably close this gap. In contrast, SEAL outperforms unprotected baselines across six models and three random seeds. Reliable self-improvement need not abandon self-verification, but it requires at least one deployment-acceptance signal outside the agent’s control.

[MA-3] Algorithms for Equilibria in Concurrent Stopping Games

【速读】:该论文旨在解决并发博弈(concurrent games)中纳什均衡(Nash equilibrium)存在性问题的不可判定性,特别是针对约束型存在性问题——即是否存在一个纳什均衡,使得每位玩家的期望收益均落在给定区间内。该问题在一般情况下是不可判定的,即使在10人停止博弈(stopping games)中依然如此。为实现可解性,论文提出两种关键策略:其一是放松精确性要求,引入ε-纳什均衡(ε-Nash equilibrium, ε-NE)框架,研究近似约束存在性问题,设计出时间复杂度为指数级但关于ε的位大小仅多项式依赖的算法;同时证明该问题在回合制博弈中已达到PSPACE-hard下界。其二是弱化解概念,采用近期提出的极端风险敏感均衡(extreme risk-sensitive equilibria, XRSE),其中玩家被划分为乐观者与悲观者,分别以正概率可达的最佳或最差收益作为评价标准,而非期望收益。论文进一步证明,在并发博弈中XRSE的约束存在性问题为NP完全,与回合制博弈情形一致,从而实现了理论上的可处理性。

链接: https://arxiv.org/abs/2607.24219
作者: Léonard Brice,Thomas A. Henzinger,K. S. Thejaswini
机构: 未知
类目: Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Concurrent games are a standard model for multi-agent systems, with Nash equilibrium as their central solution concept. The associated \emphconstrained existence problem—does a game admit a Nash equilibrium whose expected payoff lies within a prescribed interval for every player?—is undecidable, and remains so even for 10-player \emphstopping games, in which a terminal state is reached almost surely under every strategy profile. We give two routes to tractability. We first relax exactness and consider the problem of approximate constrained existence problem, parametrised by \varepsilon -NE, which decides whether an (\varepsilon)-Nash equilibrium with the prescribed payoffs exists. The algorithm runs in exponential time, and only polynomially in the bit-size of (\varepsilon). We complement it with a \PSPACE-hardness lower bound that holds already for turn-based games, and for pure equilibria as well. We then relax the solution concept, turning to \emphextreme risk-sensitive equilibria (XRSE), recently introduced for turn-based stochastic games. Here the players are partitioned into optimists and pessimists, who evaluate a strategy profile by the best, respectively the worst, payoff attainable with positive probability, instead of the expected payoff. We prove that the constrained existence problem for XRSE is \NP-complete on concurrent games, as for turn-based games. Subjects: Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA) Cite as: arXiv:2607.24219 [cs.GT] (or arXiv:2607.24219v1 [cs.GT] for this version) https://doi.org/10.48550/arXiv.2607.24219 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[MA-4] Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

【速读】:该论文旨在解决现代多智能体知识系统中知识传递链条的可信度评估问题,即如何在由多个自主转换步骤构成的知识生成与传播过程中,对每个主张(claim)的可靠性进行分级评估。现有研究虽能记录执行轨迹、工具调用和证据链接等溯源信息,并具备源可靠性估计方法(如真理发现、声誉系统),但缺乏一个可操作的框架来实现基于领域特性的、逐级传输链的发送者可信度量化,且未能涵盖完整性语义、类型化变换聚合、内容批判与路由分离等关键需求。其解决方案的关键在于借鉴古典伊斯兰圣训学(hadith science)的成熟体系,将“传系”(isnad,完整传输链)、“人物评鉴”(rijal,叙述人诚信与精确性分级)、最弱环节评估、独立链交叉验证及“文本批判”(matn criticism,内容独立于链路质量评价)等核心机制形式化映射至多智能体系统架构中。作者提出了一种从圣训科学概念到多智能体流水线的正式映射,构建了支持主张链与分级叙述人注册表的关系型数据模式,设计了一个结合链路评分与内容批判的决策矩阵,并在20,000条来自真实物理教科书的主张上进行了实证评估。结果验证了最弱环节隔离与独立链交叉验证的有效性,揭示了评分恢复环路部分失效(未能识别最高故障叙述人),并指出两项分析因数据限制而结论不明确,包括一项匹配覆盖度对比实验。全文始终明确界定哪些主张得到了证据支持,哪些尚未得到支持,体现出高度的透明性与严谨性。

链接: https://arxiv.org/abs/2607.24117
作者: Ali Zahid Raja
机构: Independent Researcher(独立研究员)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 25 pages. Code, data, and analysis: this https URL (v2.0.4). Software archived at doi: https://doi.org/10.5281/zenodo.21216873

点击查看摘要

Abstract:Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to claim-level transmission chains, with completeness semantics, transformation-typed aggregation, decoupled content criticism, and serve/review/quarantine routing. Classical Islamic hadith science confronted a structurally similar problem: deciding whether knowledge transmitted through chains of human narrators should be accepted. Over centuries it developed a rigorous methodology - isnad (a complete transmission chain attached to every claim), rijal (systematic grading of each narrator’s integrity and precision), weakest-link chain evaluation, corroboration through independent chains, and matn criticism (content evaluated independently of chain quality). This paper transfers that methodology to AI system design. We contribute a formal mapping from hadith-science concepts to multi-agent pipelines, a relational schema implementing claim chains and a graded narrator registry, a decision matrix combining chain grade with content criticism, and an evaluation on 20,000 claims from real physics textbooks. The evaluation validates weakest-link quarantine and independent-chain corroboration; reports a partial failure of the grade-recovery loop, which missed the highest-fault narrator; and reports two analyses as inconclusive, including a matched-coverage comparison the framework could not reach with the reference content critic. The paper is explicit throughout about which claims the evidence does and does not yet support. Comments: 25 pages. Code, data, and analysis: this https URL (v2.0.4). Software archived at doi:https://doi.org/10.5281/zenodo.21216873 Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA) ACMclasses: I.2.11; H.3.5 Cite as: arXiv:2607.24117 [cs.AI] (or arXiv:2607.24117v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.24117 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[MA-5] Moral Hazard in Multi-Agent Language Models

【速读】:该论文旨在解决在社会价值高但成本高昂、信息隐蔽且主要惠及他人的努力中,合作机制容易失效的问题。其核心挑战在于如何激励语言代理在面临局部短期收益与长期团队协作利益冲突时,主动披露隐藏信息以促进集体决策的优化。解决方案的关键在于构建“对话道德风险博弈”(Dialogue Moral Hazard Game),通过一个受控的文本交互框架,将隐性行动结构形式化为可测量的行为范式:在每轮交互中,语言代理需权衡是否支付查询成本以揭示对另一代理下游决策至关重要的安全事实。研究评估了七种开源大模型,并从查询使用、信息传递实效、局部奖励保留、非安全选择、格式有效性及团队成功等多个维度进行行为分解。结果表明,基础模型普遍倾向于保留局部奖励,却缺乏有效信息共享或协作行为。后续采用监督微调(SFT)、RLOO、SFT+RLOO序列训练及GEPA提示优化等诊断性更新方法,发现不同优化策略对行为的影响具有异质性——仅OLMo-7B展现出与机制一致的权重层级改进;而GEPA虽可能提升团队成功率,却常抑制或消除高成本查询行为,导致合作机制未被真正恢复。因此,该研究强调,单纯依赖团队成功率作为评价指标存在误导性,必须引入机制层面的行为分析,才能准确评估优化是否真正实现了预期的合作激励。

链接: https://arxiv.org/abs/2607.23982
作者: Dane Malenfant
机构: McGill University (麦吉尔大学); Mila - The Québec AI Institute (魁北克人工智能研究所)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注: Accepted to the Second Workshop on Social Simulation with LLMS: Fidelity in Applications at COLM 2026

点击查看摘要

Abstract:Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmström’s team moral-hazard model, we introduce the Dialogue Moral Hazard Game, a controlled textual game that operationalizes this hidden-action structure for language agents. In each episode, an agent can preserve an immediate local reward or pay a query cost to reveal a hidden safety fact that primarily helps another agent’s downstream decision. We evaluate seven open-weight language models and decompose behavior into query use, realized information transfer, local-reward preservation, unsafe choice, format validity, and team success. Base models commonly preserve local reward without team success or query without communicating information that changes the final decision. We then use supervised fine-tuning, RLOO, sequential SFT+RLOO, and GEPA prompt optimization as diagnostic update mechanisms. Their effects are heterogeneous: OLMo-7B shows the clearest mechanism-consistent weight-level improvement, whereas GEPA sometimes improves team success while reducing or eliminating costly queries. Thus, optimization can shift aggregate reward without recovering the intended cooperative mechanism, motivating evaluations that report mechanism-level behavior rather than team success alone.

[MA-6] SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception

【速读】:该论文旨在解决当前车联网(V2X)协同感知算法发展中面临的两大核心问题:一是缺乏大规模、多模态、多任务的合成数据集以支持统一空间表征(如鸟瞰图,BEV)的鲁棒算法开发;二是真实世界中同步多智能体数据的采集与标注成本过高,导致现有V2X数据集在规模和多样性上严重受限。其解决方案的关键在于提出SimBEV2X——一个基于CARLA仿真平台的先进合成数据生成工具,能够自动生成多样化的驾驶场景,自动采集包含激光雷达、摄像头等多模态传感器数据,并配套生成丰富的真值标签,包括3D边界框(含唯一轨迹ID)、高精地图、BEV语义分割图以及语义占据体素网格等。基于此,研究构建了目前最大规模的V2X感知数据集SimBEV2X,涵盖258个场景、102,200帧、超过58万帧激光点云、300多万张图像及2700多万个标注边界框。此外,论文提出了CoBEVFusion架构,通过将CoopDet3D与融合轴向注意力(Fused Axial Attention, FAX)结合,实现上下文感知的多智能体特征聚合,在基准测试中显著提升了性能。整体方案有效突破了数据瓶颈并推动了协同感知技术的发展。

链接: https://arxiv.org/abs/2607.23910
作者: Goodarz Mehr,Sepideh Gohari,Montasir Abbas,Azim Eskandarian
机构: 未知
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Robotics (cs.RO)
备注: Submitted to IEEE for review

点击查看摘要

Abstract:Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range. However, the development of robust V2X algorithms, particularly those relying on unified spatial representations like bird’s-eye view (BEV) representation, is hampered by the lack of large-scale, multi-modal, multi-task datasets. Moreover, collecting and annotating a large set of synchronized, real-world multi-agent data is prohibitively expensive. This has resulted in a landscape where existing V2X datasets are notably limited in both size and scope. To overcome this, we introduce SimBEV2X, an advanced synthetic data generation tool built on the CARLA simulator. SimBEV2X automatically creates randomized driving scenarios to collect multi-modal sensor data alongside various types of ground truth including 3D bounding boxes with unique track IDs, HD map information, BEV segmentation maps, and semantic occupancy voxel grids from both vehicles and RSUs. We also present the SimBEV2X dataset, the largest V2X perception dataset to date. The dataset comprises 258 scenes, each involving up to 8 connected vehicles and up to 4 RSUs across a variety of road networks. The SimBEV2X dataset is an order of magnitude larger than existing V2X datasets and contains 102,200 frames, 588,520 lidar point clouds, more than 3 million images, over 27 million bounding boxes, and a comprehensive set of other annotations. Finally, we establish a strong baseline on the SimBEV2X dataset using CoopDet3D and propose CoBEVFusion, a novel architecture that combines CoopDet3D with fused axial attention (FAX) for context-aware multi-agent feature aggregation, resulting in superior performance. SimBEV2X, the SimBEV2X dataset, and CoBEVFusion are available at this https URL and this https URL.

[MA-7] A Comparative Study of MCP and A2A for Inter-Agent Coordination in LLM -Based Systems

【速读】:该论文旨在解决异构、基于大语言模型(LLM)的智能体系统中多智能体协同与协议设计的实际挑战,特别是在复杂任务场景下如何有效实现跨智能体的通信与协调。其核心问题是:在受限的LLM驱动系统环境中,不同通信协议在支持多智能体协作时所表现出的实现复杂度、功能完备性与可维护性之间的权衡。解决方案的关键在于通过实证比较两种代表性协议——模型上下文协议(Model Context Protocol, MCP)与智能体到智能体协议(Agent2Agent, A2A),从多智能体系统工程视角评估其在真实软件工程任务中的表现。研究发现,MCP凭借轻量级实现模型和较低的协调复杂度,能够有效支持基础的跨智能体协作,但需在应用层显式处理对话状态管理与任务生命周期等协调问题;而A2A则在协议层面提供了对有状态、多轮交互的原生支持,通过任务与生命周期管理抽象提升了协同能力,但代价是显著增加的实现与协调复杂度。因此,该研究揭示了协议抽象如何影响协调责任的分布,并强调了在特定协作模式下两种协议的设计取舍,而非泛化地宣称其优劣,为现代代理系统的设计提供了基于经验的实践指导。

链接: https://arxiv.org/abs/2607.23884
作者: Ionut Predoaia,Tuong Manh Vu,Konstantinos Barmpis,Dimitris Kolovos,Antonio García-Domínguez
机构: 未知
类目: Multiagent Systems (cs.MA)
备注: 18 pages

点击查看摘要

Abstract:Recent industry practice has seen the rapid emergence of agentic systems composed of heterogeneous, tool- and LLM-mediated agent components, raising practical questions about inter-agent coordination and protocol design. This paper presents an implementation-grounded comparison of the Model Context Protocol (MCP) and the Agent2Agent (A2A) protocol, from a multi-agent systems engineering perspective, using an inter-agent coordination scenario involving LLM-based agents. We evaluate an MCP-based and an A2A-based multi-agent implementation of the same software engineering task against a set of requirements derived from prior literature and discussions with industry partners, including agent discoverability, multi-part messaging, multi-turn conversations, asynchronous communication, observability, interoperability, and access control. The results evidence that MCP can support inter-agent coordination in constrained LLM-based systems through a comparatively lightweight implementation model with lower coordination complexity, although coordination concerns such as conversational state management and task lifecycle handling must be implemented explicitly at the application layer. In contrast, A2A provides richer native support for stateful, multi-turn coordination through protocol-level abstractions for tasks and lifecycle management, but this comes with substantially greater implementation and coordination complexity. Given the narrow scope of the evaluated coordination pattern, these findings are presented as design observations from an empirical experience report rather than general claims of protocol suitability or superiority across broader classes of MAS, highlighting trade-offs and how protocol abstractions shape the distribution of coordination responsibilities in contemporary agentic systems.

[MA-8] MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

【速读】:该论文旨在解决智能城市空域中无人飞行器(UAV)在感知能力退化与语言指令模糊背景下,如何实现可信的多模态决策问题。现有基准普遍关注感知、导航、协作与推理等单一维度,却缺乏对物理证据、协议约束与行动风险在关键决策过程中是否仍保持耦合关系的评估。为此,研究提出MulRobBench——一个面向智慧城市环境的离线、协议条件化基准,用于评估视觉-语言-动作(VLA)UAV代理在真实运行约束下的多模态决策能力。其核心创新在于构建了一个融合真实多模态观测数据、协议级安全策略与行动级网络物理安全的统一评估框架,涵盖17个任务分类节点和12个评分维度,覆盖操作上下文理解、多模态证据仲裁、退化感知推理与风险感知行动规划四个阶段。评估采用语义评分与结构诊断相结合的方式,包含协议合规性、格式合规性、危险动作、解析失败及维度有效性等指标。实验表明,17个模型中最佳语义协议决策得分仅为0.5141,严格均值维度准确率仅0.1599,且通过20锚点模态消融实验验证了视觉与文本输入对决策的显著影响,揭示出模态信任选择、约束提取偏差、强光干扰、数据缺失及操作员简写是导致决策不稳定的主因。因此,该研究的关键解决方案在于建立一个可复现、高保真的多模态决策评估体系,推动实现具备可解释性与鲁棒性的可信无人机自主决策系统。

链接: https://arxiv.org/abs/2607.23870
作者: Belal S. Alsinglawi,Weizheng Wang,Junyi Wu,Yi Jiang,Lianhai Lin,Merouane Debbah,Izzat Alsmadi
机构: Zayed University (扎耶德大学); The University of Adelaide (阿德莱德大学); University of Emergency Management (应急管理系统大学); Khalifa University (哈利法大学); Texas A&M University–San Antonio (德克萨斯农工大学-圣安东尼奥分校)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注: 22 pages, 18 figures, 17 tables

点击查看摘要

Abstract:Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception, navigation, collaboration, and reasoning, but few assess whether physical evidence, protocol constraints, and action risk remain coupled during critical decisions. We introduce MulRobBench, an offline, protocol-conditioned benchmark for Vision-Language-Action (VLA) UAV agents in smart-city environments. MulRobBench integrates real UAV multimodal observations, protocol-level security policies, and action-level cyber-physical safety into a unified evaluation framework. The benchmark contains 3,024 samples spanning 17 task taxonomy nodes and 12 scoring dimensions across four stages: operational context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning. Evaluation combines semantic scoring with structural diagnostics, including policy compliance, format compliance, unsafe actions, parsing failures, and dimension-level validity. Across 17 multimodal models, the best semantic protocol-decision score reaches only 0.5141, while the best strict mean scoring-dimension accuracy is 0.1599. A controlled 20-anchor modality-ablation study changes 4-15 action selections per model, confirming that both visual and textual inputs influence decisions. Analysis identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability. MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints.

[MA-9] he Illusion of Secure LLM Code: Closing the Security Gap via Iterative Reprompting

【速读】:该论文旨在解决当前大型语言模型(Large Language Models, LLMs)在软件开发中自动生成安全认证代码的可靠性问题,尤其关注其在实际应用中是否能够满足安全要求。研究发现,尽管主流AI编码助手在功能实现上表现良好,但基于常规或泛化安全提示生成的代码普遍缺失关键安全防护机制,如防暴力破解、会话管理及强密码处理等。其解决方案的关键在于突破传统的单次提示(single-shot prompting)模式,提出通过迭代式“再提示”(Reprompting)策略,迫使模型进入上下文自我审计循环,从而构建具备纵深防御(defense-in-depth)特性的安全架构。实证结果表明,仅依赖明确的NIST标准提示仍不足以确保全面合规,唯有持续、标准化的验证流程才能实现真正可信赖的安全代码生成,因此论文主张企业级部署必须从一次性提示工程转向持续、符合标准的验证管道。

链接: https://arxiv.org/abs/2607.23710
作者: Ishpuneet Singh,Shreyas Mahajan,Gurjot Singh,Maninder Singh
机构: 未知
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly integrated into software development workflows, yet their ability to autonomously generate secure authentication code remains uncertain. This paper evaluates the security architecture of authentication systems generated by five prominent AI coding assistants through a bi-modal assessment framework combining static code analysis and dynamic penetration testing, mapped to NIST SP 800-63B guidelines. The study examines model behavior across four prompting strategies Basic, Secure, NIST-Based, and Reprompting to reflect varying levels of developer guidance. Empirical results demonstrate that code generated from functional or generically secure prompts consistently omits critical protections, particularly concerning brute-force resistance, session management, and robust password handling. While providing explicit, single-shot NIST context significantly improves compliance, the findings reveal that this remains structurally inadequate. Instead, iterative Reprompting: forcing models into a contextual self-auditing loop is strictly required to achieve a comprehensive, defense-in-depth security architecture. Ultimately, this study proves that current AI coding assistants do not produce secure-by-default applications, dictating that enterprise deployments must transition from single-shot prompt engineering to continuous, standards-driven verification pipelines.

[MA-10] Separating Capability from Permission: A Governance Framework for Agent ic AI Autonomy Levels

【速读】:该论文旨在解决当前在生成式 AI(Generative AI)系统中,技术能力与实际授权权限之间存在的混淆问题,即在讨论自主性时,常将系统的技术可实现能力等同于其应被允许的行为范围。其核心解决方案在于提出一个治理框架,明确区分“允许的自主性水平”(Allowed Autonomy Levels, AAL)与“自主能力水平”(Autonomous Capability Levels, ACL):前者反映在风险、监督与问责考量下系统被授权行使的自主程度,后者则刻画系统固有的技术能力。该框架构建了从反应式执行到委托操作权限的五级自主性层级,并系统分析了随着自主性提升,控制性、可逆性与问责机制的变化规律。为实现该框架的落地应用,研究提出了基于风险感知的自主性授权决策流程,阐明了风险与问责在不同层级间的演变关系,并通过一个企业级数据工程代理的实际部署案例,验证了即使系统具备高技术能力,也可根据风险容忍度、可逆性要求及组织准备度被主动约束至较低的允许自主性水平。这一“授权”与“能力”的解耦设计,为生成式 AI 系统的设计、部署与治理提供了可操作的指导原则。

链接: https://arxiv.org/abs/2607.23438
作者: Haining Zheng,Qian Dong,Rodolfo K. Depena,Jonathan D. Bhatia,Feng Xiao,Peng Xu
机构: 未知
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Multiagent Systems (cs.MA)
备注: 10 pages, 2 tables, 3 figures

点击查看摘要

Abstract:As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing with what they should be permitted to do in practice. This paper introduces a governance framework that explicitly separates Allowed Autonomy Levels (AAL), which define the degree of autonomy an AI agent is authorized to exercise given risk, oversight, and accountability considerations, from Autonomous Capability Levels (ACL), which characterize an agent’s inherent technical abilities. We present a structured set of autonomy levels spanning reactive execution, decision support, supervised action, goal-directed autonomy, and delegated operational authority, and describe how control, reversibility, and accountability change as autonomy increases. To operationalize this framework, we propose a risk-aware decision process for assigning allowed autonomy, analyze how risk and accountability evolve across autonomy levels, and demonstrate its application through a deployed enterprise data engineering agent, illustrating how a system assessed at a high capability level can be deliberately constrained to a lower allowed autonomy based on risk, reversibility, and organizational readiness. By distinguishing authorization from capability, this work provides practical guidance for the design, deployment, and governance of Agentic AI systems.

[MA-11] Emergent Behaviour in Financial Markets

【速读】:该论文旨在解决复杂系统中涌现行为(emergent behaviour)的自动化分析难题,尤其聚焦于电子金融市场的建模与推理。其核心问题在于,复杂系统由基本代理间的交互产生非线性、不可预测的整体行为,而传统形式化方法在处理此类系统时面临显著挑战。解决方案的关键在于不依赖特定技术框架或简化假设,而是系统性地识别并结构化复杂性的来源,包括异质性代理、动态交互、环境反馈等,并在此基础上提出可保留领域实际语义的替代性技术路径。通过强调对市场机制的形式化规范与分析的系统性研究框架,论文为实现对复杂系统中涌现现象的可靠计算支持提供了理论基础与方法论指导。

链接: https://arxiv.org/abs/2607.23311
作者: Omar Inverso,Emilio Tuosto,Dragisa Zunic
机构: 未知
类目: Multiagent Systems (cs.MA)
备注: Accepted at ISoLA 2026, Resourceful Engineering of Complex and Autonomous Systems (ReoCAS) track

点击查看摘要

Abstract:Some properties of so-called complex or collective systems can be observed to emerge from the interactions of elementary agents. This phenomenon, known as emergent behaviour, has long since been studied in the most diverse disciplines, with recent growing awareness from the formal methods community about the opportunity of opening up to seemingly distant disciplines with appropriate technology for computer-aided reasoning. Different peculiar elements of complexity make automated reasoning on these systems particularly challenging. We consider electronic financial markets to drive our discussion. We identify and structure the sources of complexity to tackle in order to provide computational support for the analysis of emergent phenomena. We refrain from evaluating the suitability of specific technical solutions or frameworks of preference, which would as usual require simplifying assumptions and divert from the actual phenomenon of interest. Rather, we elaborate on possible alternatives to handle some of the main technical aspects involved in automated analysis, while retaining a solid and concrete interpretation of the domain, and in doing so outline a more systematic research program for the formal specification and analysis of market mechanisms.

[MA-12] Online Fair Division with Budget Constraints

【速读】:该论文旨在解决在广义分配预算约束下,具有在线特性(goods arrive sequentially)的离散公平分配问题。核心挑战在于:物品需即时、不可逆地分配给可行的接收方或慈善机构(用于存放未分配物品),而公平性仅针对每个接收方包中预算可行的子集进行评估。由于缺乏额外结构时,任何确定性在线算法均无法保证对预算可行的无嫉妒性(feasible envy-freeness)有固定近似比,即使在高度对称实例中亦然。为克服此局限,论文提出“有界密度分布”(bounded density spread)作为关键结构性条件,由此可获得对任意物品大小的近似算法;在常见估值函数及小物品假设下,该近似可达到最优确定性边界。此外,研究引入资源增强(resource augmentation)机制,即允许在线算法使用略大于公平基准的预算,从而刻画其对可实现公平性保障的提升效果。最后,论文构建了一种基于联合价值-大小类型预测的强化学习框架,证明了在完美预测下的一致性与对预测误差的鲁棒性,同时指出仅独立预测价值和大小边缘分布不足以恢复强公平性保障。

链接: https://arxiv.org/abs/2607.23310
作者: Saar Cohen,Nicholas Teh,Paul W. Goldberg,Michael J. Wooldridge
机构: 未知
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Theoretical Economics (econ.TH)
备注:

点击查看摘要

Abstract:We study an online variant of discrete fair division under generalized assignment budget constraints. Goods arrive one at a time and must be assigned irrevocably to a feasible agent or to charity, which holds all unallocated goods, while fairness is evaluated only against budget-feasible subsets of every recipient’s bundle. We first show that, without additional structure, no deterministic online algorithm can guarantee any fixed approximation to feasible envy-freeness, even in highly symmetric instances. We then identify bounded density spread as a structural condition that restores meaningful guarantees, obtaining approximation algorithms for arbitrary item sizes and showing that, under common valuations and sufficiently small goods, these guarantees can be strengthened to an optimal deterministic frontier. We further study resource augmentation, where the online algorithm is allowed slightly larger budgets than the fairness benchmark, and characterize the resulting improvement in the achievable guarantees. Finally, we develop a learning-augmented framework based on predicting joint value-size types, proving consistency under perfect predictions, robustness to prediction error, and showing that separate predictions of value and size marginals are insufficient to recover strong fairness guarantees.

[MA-13] Co-Evolving Graph and Text Memory for Training-Free Multi-Hop Question Answering

【速读】:该论文旨在解决多跳问答(multi-hop question answering)中跨推理步骤协调关系型证据与文本证据的难题,这一需求仅依赖纯文本语料库或知识图谱(Knowledge Graph, KG)均无法单独满足。现有方法通常仅关注该过程的部分环节:基于图的检索增强生成(Graph-augmented RAG)从预构建或查询更新的图中检索信息,知识图谱问答(KGQA)系统在主题中心子图中搜索,而记忆增强型智能体则维护动态记忆但未持续将图谱记忆与文本上下文对齐。为应对上述局限,本文提出Co-E,一种无需训练的系统,其核心在于构建同步双向的图-文工作记忆机制。该系统的同步周期实现文本记忆的固化、关系三元组的提取并注入图谱记忆,同时将图谱事实回注至生成上下文,使两种记忆持续相互影响并指导后续的检索与生成。实验在六个多跳问答基准上验证了Co-E的有效性,结果表明其在不依赖训练的开放主干模型中表现优于现有基线,并可与更大规模或需训练的系统相媲美。

链接: https://arxiv.org/abs/2607.23278
作者: Hieu Man,Thien Huu Nguyen
机构: University of Oregon, OR, USA
类目: Computation and Language (cs.CL); Multiagent Systems (cs.MA)
备注:

点击查看摘要

Abstract:Multi-hop question answering requires coordinating relational and textual evidence across reasoning steps, a combination neither a text corpus nor a knowledge graph can supply alone. Prior work often emphasizes only part of this loop: graph-augmented RAG retrieves from a pre-built or query-updated graph, KGQA systems search within topic-centered subgraphs, and memory-augmented agents maintain evolving memories without continuously reconciling graph memory with textual context. We propose Co-E, a training-free system built around synchronized bidirectional graph-text working memory. A synchronization cycle consolidates textual memory, extracts relational triples into graph memory, and injects graph facts back into the generation context. Because both memories are maintained, they shape subsequent retrieval and generation. Evaluated on six multi-hop QA benchmarks, Co-E improves over comparable training-free open-backbone baselines and is competitive with larger or trained systems.

[MA-14] Let AI Agents Translate Networks Not Reason About Them

【速读】:该论文旨在解决生产环境中缺乏可信赖的网络形式化模型(formal model)这一核心问题,其根本挑战在于手工构建和维护此类模型需要罕见的专业知识,且难以随网络频繁变更而保持同步。论文提出的关键解决方案是将网络建模过程从人工主导转向由大语言模型(LLM)驱动的自动化符号翻译——即将网络配置、拓扑及路由状态等实际运维数据转化为形式逻辑规则,这一任务恰好是当前大语言模型在结构化输出方面具备优势的领域。与依赖生成式AI进行端到端自主推理的主流趋势相反,本文主张将AI限制于“翻译”环节,并借助形式化求解器(solver)实现长期、可靠的推理能力。由此构建的可复用、可验证的形式化模型,能够被灵活应用于具体运维任务如根因分析(RCA)、可达性验证与变更影响分析等。作者提出的TypoNet系统通过在模拟的生产级广域网(WAN)环境中自动生成并验证符号模型,初步评估表明,该方法不仅独立执行时比直接使用LLM更高效、低成本且可靠,作为辅助工具还可显著降低故障定位的成本。因此,该研究的核心贡献在于确立了一种以可验证建模为基础、结合求解器保障推理可靠性的人工智能应用范式,为复杂网络系统的自动化运维提供了可信的技术路径。

链接: https://arxiv.org/abs/2607.22947
作者: Hongyu Hè,Maria Apostolaki
机构: Princeton University (普林斯顿大学)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Networking and Internet Architecture (cs.NI); Symbolic Computation (cs.SC)
备注: 8 pages, 3 figures, 1 table

点击查看摘要

Abstract:A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no production network has one, since writing a model by hand demands rare expertise and is hard to keep current as the network changes frequently. At its core, network modeling is a typographical exercise: it translates network artifacts (e.g., configurations, topology, and routing state) into rules in formal logic. Translation of this kind is what large language models (LLMs) nowadays do well. Unlike free-form AI reasoning, such translation can be formally verified. Once modeling is no longer the bottleneck, trusting AI to reason over large, complex networks no longer makes sense. Our position therefore cuts against the prevailing race to put autonomous AI agents in charge end-to-end. We instead confine AI to translation and rely on a solver for reliable long-horizon reasoning, building a reusable formal model of general network behavior that can then be specialized to specific tasks, e.g., root-cause analysis (RCA). We build TypoNet that constructs and validates a symbolic model of an emulated production-scale WAN from the network’s own artifacts. Our preliminary evaluation shows TypoNet helps in two ways. On its own, TypoNet answers operational questions (e.g., reachability verification and change-impact analysis) faster, more cheaply, and more reliably than an LLM. As a tool for an AI agent, TypoNet boosts fault localization at lower cost. The result makes the case for AI that builds verifiable network models and relies on a solver for reliable long-horizon reasoning. Comments: 8 pages, 3 figures, 1 table Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Networking and Internet Architecture (cs.NI); Symbolic Computation (cs.SC) MSC classes: 68M15 (Primary), 68Q60, 68T27, 68T30 (Secondary) ACMclasses: C.2.3; D.2.4; I.2.3 Cite as: arXiv:2607.22947 [cs.AI] (or arXiv:2607.22947v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.22947 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[MA-15] Agent Team Work Zone: An Automated Persistent Workspace for Long-Lived Coding Agent Teams

【速读】:该论文旨在解决大型语言模型(Large Language Model, LLM)编码代理在长期任务执行中面临的四大核心问题:不可恢复的代理团队状态、信息压缩导致的工作细节丢失、代理操作积累的技术债务,以及繁琐的提示词编写。其解决方案的关键在于提出一种基于文件系统的操作层——ATWZ(Agent Team Work Zone),该系统以“每个代理及其成员如同人类员工”为核心设计原则,将代理的关键工作状态持久化存储于专用目录“工作站”(workstation)中的文件里,并与相应的技能、钩子(hooks)和脚本协同管理。通过这一机制,代理团队可定期备份工作状态,在进程中断后仅需一条命令即可恢复,有效避免了因压缩(compaction)造成的信息丢失;同时,通过文件化协作与文档传递,显著降低了重复编写复杂提示词的需求,并大幅缓解了长期积累的代理“技术债务”问题。

链接: https://arxiv.org/abs/2607.22917
作者: Shouren Wang
机构: 未知
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: 31 pages, 9 figures

点击查看摘要

Abstract:Large Language Model (LLM) agents have significantly improved coding and programming workflows. Claude Code, in particular, is one of the most powerful LLM coding agents and is capable of conducting complex coding tasks. However, several drawbacks can undermine long-term agentic workflows. (1) Irrecoverable agent teams: The Agent Teams feature is powerful, but the working state accumulated by each teammate is lost and cannot be resumed once the process stops, for example, when a terminal is closed. (2) Compaction erodes working detail: Compaction condenses the conversation into a summary, causing an agent’s working details to become vague. (3) Agentic “technical debt”: Over time, a user’s decisions and the agents’ operations become trapped in compacted old chats, making the project increasingly difficult to maintain and review. (4) Heavy prompt writing: Assigning or handing off tasks requires users to repeatedly write long prompts to achieve the expected agentic performance. We propose ATWZ (Agent Team Work Zone), a filesystem-based operations layer built around Claude Code’s native Agent Teams that addresses these problems. Its central design principle is to treat each agent and teammate as a human employee and preserve their important working state in files stored in a dedicated directory called a “workstation,” together with the skills, hooks, and scripts that use and maintain these files. With ATWZ, an agent team can periodically back up its working state, allowing an agent’s knowledge to be recovered after compaction. After a process ends, the team can be restored with a single command. These features also substantially mitigate the agentic “technical debt” described above. Moreover, within ATWZ, agent “employees” can send documents to one another, greatly reducing the effort required to write prompts.

[MA-16] Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks

【速读】:该论文旨在解决多智能体诊断框架中迭代大语言模型(LLM)因通信拓扑结构不合理而导致的诊断安全与可靠性问题。现有架构普遍采用无标度(scale-free)或小世界(small-world)网络,假设其具备最优通信效率,但本研究通过在768维Bio_ClinicalBERT语义嵌入空间中,利用巴拉萨-阿尔伯特(Barabási–Albert, BA)和瓦茨-斯特罗加茨(Watts–Strogatz, WS)网络构建解析的各向同性方差代理,数学上证伪了该假设在语义数据场景下的有效性。其关键发现是:局部高密度团簇(dense cliques)会引发结构性瓶颈,导致幻觉数据被局部限制,无法实现全局共识,系统最终陷入熵饱和阈值 $ H_\infty \approx 5.947 ,并伴随严重的终端余弦相似度衰减(53.29,并伴随严重的终端余弦相似度衰减(53.29%),完全覆盖原始真实语义。此外,高度聚类架构表现出灾难性的方差放大(51.81%,\rho = 1.5181),远超随机图(Erdo\HsReˊnyi),远超随机图(Erdős–Rényi,\rho = 1.0766$),表明系统完全不可预测。因此,该研究提出一种基于图拉普拉斯矩阵连续特征分解的动态谱监控机制,以 O(N3)\mathcal{O}(N^3) 时间复杂度确保代数连通性(algebraic connectivity, λ2\lambda_2)的严格下界,从而实现全局状态扩散,将拓扑稳定性作为保障自主医疗诊断可靠性的量化刚性要求。

链接: https://arxiv.org/abs/2607.22758
作者: Amritesh Banerjee
机构: University of Massachusetts Amherst (马萨诸塞大学阿默斯特分校)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying communication topologies. Frequently used architectural paradigms depend on scale-free or small-world networks, assuming optimal communication efficiency. Our study mathematically dismantles that assumption for semantic data. By mapping multi-agent communication uncertainty trajectories onto a 768-dimensional Bio_ClinicalBERT embedding space via an analytical isotropic variance proxy using Barab’asi–Albert (BA) and Watts–Strogatz (WS) networks, we prove that structural bottlenecks compromise diagnostic safety. Our phase transition matrices illustrate that localized dense cliques confine hallucinated data, preventing global consensus and forcing the system toward a permanent entropy saturation threshold of H_\infty \approx 5.947 . As a result, we measure a severe terminal cosine similarity degradation of 53.29%, completely overwriting the original ground-truth. Moreover, the terminal semantic drift reveals a catastrophic variance amplification of 51.81% ( \rho = 1.5181 ) in highly clustered architectures, proving total system unpredictability when compared to Erdős–R’enyi configurations ( \rho = 1.0766 ). Instead of reducing errors, hub-centric systems autonomously compound localized hallucinations. By introducing dynamic spectral monitoring operating at an \mathcalO(N^3) time complexity and imposing a strict lower bound on algebraic connectivity ( \lambda_2_min ) via the continuous eigen-decomposition of the graph Laplacian, we present a mathematically rigorous technique to ensure global state diffusion. Securing the reliability of autonomous medical diagnostics necessitates treating topological stability as a non-negotiable quantitative imperative.

[MA-17] A Vocabulary for Multi-Agent Automated Research Systems

【速读】:该论文旨在解决自动化研究系统(autoresearch systems)在多智能体架构设计中缺乏统一描述与比较框架的问题,尤其针对不同系统在智能体角色、操作能力、通信机制、信息持久性、行动决策、运行初始化及输出评估等方面的异构性带来的可比性难题。其核心解决方案是提出一套系统化的词汇表(vocabulary),用于精确刻画多智能体自动化研究系统的关键组成部分与动态行为,包括:1)智能体身份与职责;2)可用操作及其调用权限;3)智能体间通信方式;4)系统内与跨运行的信息可见性;5)下一步行动的选择机制;6)运行启动方式;7)输出评估策略。该词汇表将原本模糊的结构设计问题转化为可测试的具体选择,同时明确将评估器作为系统的一部分,从而区分“生成式品味”(generative taste,即系统在未获得评分前提出新颖轨迹的能力)与“评价式品味”(evaluative taste,即代理评分与真实质量之间的差距),使对系统“缺乏品味”的批评得以分解为两个可分别优化的技术问题。通过在近期自研究系统上的实例化应用,验证了该词汇表能够有效覆盖结构差异显著的系统设计。

链接: https://arxiv.org/abs/2607.22682
作者: Bardiya Akhbari
机构: Amazon AGI
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: 22 pages, 4 figures, 3 tables

点击查看摘要

Abstract:We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The vocabulary specifies 1) who the agents are, 2) what operations are available in the system, 3) who may invoke them, 4) how agents communicate, 5) what information is visible within and across runs, 6) how the next action is chosen, 7) how a run begins, and 8) how outputs are evaluated. A trajectory records one run from the input task to the returned artifact. Because agents, operations, and initialization may be stochastic, repeated runs on the same task induce a distribution over trajectories rather than a single behavior. Our vocabulary turns structural design questions, such as when agents should communicate, gain or lose a capability, or carry information across runs, into testable choices. It also makes the evaluator a component of the system, since reported gains depend on how closely the proxy score matches true quality. That separation also splits the vague complaint that these systems lack taste into two failures with different solutions. Generative taste is the rate at which a system proposes novel trajectories before any score is observed, and evaluative taste is the gap between the proxy score and the quality it should match. We instantiate the vocabulary on recent autoresearch systems to illustrate that it covers designs that differ widely in structure. Comments: 22 pages, 4 figures, 3 tables Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA) ACMclasses: I.2.11; I.2.7; I.2.6 Cite as: arXiv:2607.22682 [cs.AI] (or arXiv:2607.22682v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.22682 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[MA-18] CRAFT: Learn the Schema Execute the Plan

【速读】:该论文旨在解决企业级生成式编码代理(coding agent)在处理自然语言分析请求时,因在提示词中注入详尽的模式(schema)与工具文档而导致的推理开销大、模式演化复杂及多轮分析可靠性下降的问题。其核心挑战在于如何在不依赖实时提示注入完整模式信息的前提下,实现对复杂企业数据环境中的模式知识与工具使用行为的稳定掌握。解决方案的关键在于提出一种两阶段后训练框架CRAFT:第一阶段采用去模式化计划监督微调(schema-stripped PLAN SFT),从经过验证的执行轨迹中学习领域结构化的计划与可执行行为,避免了运行时对完整模式的依赖;第二阶段通过执行导向的强化学习(execution-shaped reinforcement learning),优化工具选择策略、代码质量、计划-代码一致性以及失败执行的恢复能力。训练轨迹通过三重门控过滤器(Tri-Gate filter)进行筛选,结合执行验证、数据完整性检查与大模型判断审计,确保高质量学习信号。实验表明,相较于传统“模式堆叠”基线,CRAFT在广告分析场景下显著提升综合代理得分(+9.6个百分点)、一致性(+4.1个百分点)和多轮连贯性(+4.2个百分点),同时将输入令牌消耗降低约9倍,模式发现循环减少达5倍,有效支撑了生产环境下的可扩展、高可靠分析代理部署。

链接: https://arxiv.org/abs/2607.22642
作者: Aakash Kolekar,Sahika Genc,Shahriar Shariat,Bunyamin Sisman,Tibor Mezi,Barbara Poblete,Shree Vandana Kachroo,Calvin Chi,Parth Parmar,Ari Singer,Prayaas Jain,Cindy Barker,Benoit Dumoulin
机构: Amazon Advertising Foundations(亚马逊广告基础架构); Amazon Web Services Agentic AI(亚马逊网络服务代理人工智能)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric definitions. Yet the prevailing deployment pattern injecting exhaustive schema and tool documentation into each prompt increases inference overhead, complicates schema evolution, and undermines reliability in multi-turn analysis. We investigate whether stable schema knowledge and tool-use behavior can instead be acquired through post-training while preserving the consistency required for production-facing analytics. We present CRAFT, a two-stage post-training recipe for schema-grounded coding agents. First, schema-stripped PLAN supervised fine-tuning learns domain-structured plans and executable behaviors from validated trajectories without exhaustive prompt-time schema injection. Second, execution-shaped reinforcement learning aligns the policy for tool selection, code quality, plan-code consistency, and recovery from failed executions. Training trajectories are curated through a Tri-Gate filter combining execution validation, data-integrity checks, and LLM-judge reasoning audit. We evaluate CRAFT for planned rollout in advertising analytics, covering campaign performance analysis, metric drill-downs, entity-level performance analysis, and multi-turn analytical refinement. The enterprise evaluation environment incorporates beta APIs as the agent-facing tool surface and spans 25 schema-linked core entities and 30 agentic workflows. Relative to a schema-stuffed baseline, CRAFT improves composite Agent Score by +9.6 pp, consistency by +4.1 pp, and multi-turn coherence by +4.2 pp, while reducing input-token burden by approximately 9x and schema-discovery loops by up to 5x. We further report deployment tradeoffs, reward-shaping limitations, and training-infrastructure extensions required for multi-turn tool-use reinforcement learning in enterprise settings.

[MA-19] Decentralized Granular Access Control for Agent ic AI Systems in Critical Infrastructure

【速读】:该论文旨在解决在生产环境中部署自主型AI代理(autonomous AI agents)所面临的根本性安全挑战,特别是传统基于角色的访问控制(RBAC)模型无法有效应对非确定性行为带来的风险。由于AI代理具有随机性(stochastic behavior),其行为难以预测,导致依赖静态权限的现有信任模型失效。为此,论文提出一种专为关键云基础设施中运行的智能体系统设计的去中心化、多层访问控制架构。其解决方案的关键在于四项核心创新:(1)复合身份模型,将代理行为与委托的人类权威绑定;(2)五级细粒度的分层权限体系,涵盖从全局平台访问到单个参数级别的约束;(3)去中心化的策略所有权机制,使各工具团队可独立管理自身授权边界;(4)具备安全锁机制的渐进式信任升级,防止自主代理执行高风险操作。该架构基于OWASP LLM应用十大威胁(2025版)威胁分类体系进行设计,每一项决策均针对特定攻击向量提供防御。实际部署于大型云服务商,支撑超过20个专用AI代理及60多个确定性剧本,日均处理数千次操作,连续八个月零未经授权写操作记录。实证数据表明,该多层授权机制在抑制非确定性主体的权限提升方面具有显著有效性。

链接: https://arxiv.org/abs/2607.22611
作者: Arun Malik,Deepal Jayasinghe,Bradley Klemick,Prachi Shah,Nitish Talasu,Vineet Tushar Trivedi
机构: 未知
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Emerging Technologies (cs.ET); Multiagent Systems (cs.MA)
备注: 7 pages, 9 figures, 7 tables

点击查看摘要

Abstract:The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based access control (RBAC) models cannot address. Unlike deterministic automation, AI agents exhibit stochastic behavior, making conventional trust models insufficient for governing their access to critical systems. This paper presents a decentralized, multi-layered access control architecture designed specifically for agentic AI systems operating in critical cloud infrastructure. Our framework introduces four key innovations: (1) a compound identity model that binds agent actions to delegated human authority, (2) a hierarchical permission system spanning five granularity levels from global platform access to per-parameter constraints, (3) a decentralized policy ownership model where tool teams independently govern their authorization boundaries, and (4) progressive trust escalation with safety interlocks that prevent autonomous agents from executing high-risk operations. We ground our design in the OWASP Top 10 for LLM Applications (2025) threat taxonomy and demonstrate how each architectural decision mitigates specific attack vectors. Deployed in production at a major cloud provider managing network infrastructure across hundreds of datacenters, the system enforces granular access control for 20+ specialized AI agents and 60+ deterministic playbooks processing thousands of operations daily while maintaining zero unauthorized write operations over eight months of production deployment. We present empirical data on access pattern distributions, denial rates, and the effectiveness of layered authorization in preventing privilege escalation by non-deterministic actors.

[MA-20] Lexical discovery in unknown environments orchestrated by Large Language Models

【速读】:该论文旨在解决自主代理群体在未知环境中(如行星或深海探测)缺乏共同语言标签,无法有效交流感知实体的问题。其核心挑战在于如何在无先验人类语言命名的情况下,实现多代理系统自主构建共享的、可理解的“外星词汇”(alien lexicon)。解决方案的关键在于提出神经符号词汇发现(Neuro-Symbolic Lexical Discovery, NSLD)框架:每个代理结合冻结的CLIP视觉编码器与私有FAISS向量索引,利用仅文本的大型语言模型(LLM)进行语义推理;通过在分布外(out-of-distribution)视觉参照物上开展指称游戏(referential game),代理间基于嵌入空间中的语义相似性达成共识,使新发现的词汇通过语义邻近性锚定于自然语言,从而扩展人类词汇体系。该方法在模拟中成功实现了最多20个代理与10个视觉参照物之间的词汇收敛,且三种分析模型对收敛动态的拟合达到R² > 0.95,为自主探索任务的预部署规划提供了首个量化基础。

链接: https://arxiv.org/abs/2607.22591
作者: Rafael Sendra-Arranz,Iñaki Dellibarda Varela,Eduardo Rocon,Álvaro Gutiérrez,Manuel Cebrian
机构: CSIC(西班牙国家研究理事会)
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 32 pages, 7 figures

点击查看摘要

Abstract:Populations of autonomous agents deployed in unknown environments (e.g. planetary or deep-sea exploration) must develop shared vocabularies to refer to entities that have no name in any human language. We propose the Neuro-Symbolic Lexical Discovery (NSLD) framework, in which a population of LLM-based agents plays a referential game over out-of-distribution visual referents, autonomously self-organising a shared alien lexicon. Each agent combines a frozen CLIP vision encoder with a private FAISS vector index and a text-only LLM. Crucially, discovered alien words are anchored to natural language via semantic proximity in the embedding space, enlarging the human vocabulary with new perceptually grounded words. Consensus is reached in simulations with populations of up to twenty agents and ten visual referents. Convergence dynamics are characterised through three analytical models achieving R^2 0.95, representing a first step towards pre-deployment planning in autonomous exploration missions.

自然语言处理

[NLP-0] ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

【速读】: 该论文旨在解决多模态大语言模型(Multimodal Large Language Models, MLLMs)在医疗领域部署中的核心挑战:如何实现对异构二维(2D)与三维(3D)医学影像的统一理解,以及如何构建符合放射科医生临床实践、具备细粒度与事实性驱动特性的评估体系。其解决方案的关键在于提出一种以视觉为中心的架构设计——ClinFusion,采用分层级联的空间感知局部融合(Cascade Spatial-Aware Locality Fusion)视觉编码器结构,有效整合2D与原生3D医学图像的特征表示,实现跨模态的深度融合。同时,构建了基于视觉引导的评估框架,包括面向指令遵循能力的MedIF-Bench基准和基于感兴趣区域(Region-of-Interest, RoI)的报告生成评估方法,确保评估结果与临床实际高度对齐。实验表明,ClinFusion在涵盖视觉问答、报告生成与指令跟随等任务的多模态医疗基准上均达到新基准水平,在24项评测中优于20项领先开源模型,并在16项多模态任务中超越GPT-5.2与Gemini-3-Flash等先进闭源模型;此外,由认证放射科医师进行的盲评验证了其生成报告质量最高,且RoI基评估指标与专家判断的相关性最强,显著优于现有自动评估方法。

链接: https://arxiv.org/abs/2607.24743
作者: Hangjie Yuan,Yichen Qian,Zhiwei Tang,Xianzhe Xu,Lirong Wu,Sicheng Yang,Jinwang Wang,Pengju Wang,Zhitao Zeng,Yizeng Han,Yan Xing,Shengxuan Luo,Tao Feng,Qing Xie,Weigen Yao,Yi Yang,Zuozhu Liu,Jiasheng Tang,Shaocheng Wang,Jitao Wang,Jiahong Dong,Weihua Chen,Feng Xu,Fan Wang
机构: 未知
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Code: this https URL Models: this https URL

点击查看摘要

Abstract:Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists’ clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks—spanning visual question answering, report generation, and instruction following—as well as textual medical tasks, outperforming leading open-source medical MLLMs (\textite.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.

[NLP-1] he Physics of Multi-Turn Long-Horizon Planning : From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agent ic Distillation

【速读】: 该论文旨在解决基础模型智能体在多轮长时程规划(multi-turn long-horizon planning)能力上的根本性提升问题。现有模型依赖于不可控且黑箱的互联网数据进行训练,导致难以厘清规划能力的获取、塑造与整合机制。为此,本文提出一个统一且可控的多轮环境,实现对长时程规划在三个阶段的系统性研究:(1)预训练阶段的规划能力获取,通过分析数据格式、分布与质量发现,显式构建基于思维链(CoT)状态转移建模的世界模型能显著增强长时程泛化能力;原子技能不足以支撑组合泛化,而少量长时程数据即可有效提升性能,且次优轨迹会因误差累积严重损害整体表现;(2)后训练阶段的规划能力塑造,采用GRPO与OPD方法,通过互信息区分通用规划模式与任务特定规划知识,结果显示在低质量与长时程设定下,OPD因提供更一致的更新方向,其有效应用范围优于GRPO;而从不同知识结构的教师中蒸馏未见流程可能破坏学生模型原有的世界建模先验,且无法完全建立新知识;(3)后训练阶段的规划能力整合,引入多教师在线策略蒸馏(MOPD),通过收敛至跨环境共享的规划模式实现能力融合,兼容的规划模式支持跨环境泛化,部分共享模式有利于持续学习,而完全冲突的模式则引发严重干扰。其核心解决方案在于构建可控实验环境并分阶段揭示规划能力的形成机制,关键在于通过显式世界模型构建、基于互信息的规划知识解耦以及多教师协同蒸馏来实现规划能力的高效获取、精准塑造与鲁棒整合。

链接: https://arxiv.org/abs/2607.24720
作者: Tianyi Men,Zhuoran Jin,Kang Liu,Jun Zhao
机构: Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning ability acquisition during pre-training. We study data format, distribution, and quality. Explicit world model construction through CoT state transition modeling yields stronger long-horizon generalization. Atomic skills alone are insufficient for compositional generalization, whereas a litte long-horizon data works. Moreover, suboptimal trajectories severely impair performance because errors amplify over long horizons. (2) Planning ability shaping via GRPO and OPD post-training. Through mutual information, we distinguish general planning patterns from task-specific planning knowledge. For planning patterns, we identify three application regions of post-training: unnecessary, effective, and unsupported. OPD has a broader effective region than GRPO under low-quality and long-horizon settings, as it provides more consistent update directions. For planning knowledge, distilling unseen procedures from a teacher with different knowledge may impair student’s prior world modeling without fully establishing new knowledge. (3) Planning ability integration through MOPD post-training. We show that multi-teacher on-policy distillation (MOPD) integrates capabilities by converging to shared planning-pattern across environments. Compatible patterns enable cross-environment generalization, partially shared patterns support continual learning, while completely conflicting patterns cause severe interference.

[NLP-2] DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)预训练数据处理过程中存在的“一刀切”式处理策略问题,即现有方法通常在语料库或领域层面设定固定的处理规则,并对所有样本统一应用,缺乏针对每个数据样本个性化需求的自适应能力。其核心解决方案是提出DataOrchestra框架,通过构建一种示例级(example-specific)的数据处理流水线,实现精细化、动态化的数据处理。该框架的核心创新在于引入一个协调器(orchestrator),能够根据每个数据块的特性智能决策:是否丢弃、保持原样或进行清洗;对于需清洗的数据,协调器会从程序化编辑到基于大模型的重写等多种操作中选择合适的下游操作,并为每一步生成具体的指令,由相应的工具模型执行。该方法不仅在11个基准测试中均表现出稳定的性能提升,优于单一处理方法,还在数学持续预训练任务中显著超越强基线,同时通过跳过不必要的处理步骤有效降低了计算开销。

链接: https://arxiv.org/abs/2607.24717
作者: Zhen Huang,Yikun Wang,Shijie Xia,Pengfei Liu
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 36 pages

点击查看摘要

Abstract:Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without adapting to the needs of each example. We propose DataOrchestra, a framework that unifies different processing operations and orchestrates an example-specific pipeline for each example. Given a chunk of pretraining data, an orchestrator decides whether to drop, untouch, or clean it. For a chunk to be cleaned, it selects one or more downstream operations, ranging from programmatic editing to different forms of LLM-based rewriting. For each rewriting step, it further generates a concrete instruction, which is executed by the corresponding downstream tool model. We pretrain models from 0.5B to 7B from scratch on web data processed by DataOrchestra and observe stable average gains over individual data-processing methods across 11 benchmarks. DataOrchestra is also effective for math continued pretraining and outperforms stronger processing baselines, while reducing processing compute by skipping unnecessary downstream operations.

[NLP-3] Beyond Scale and Generation: Understanding Language Model-based Entity Matching

【速读】: 该论文旨在解决实体匹配任务中因模型架构、模型骨干、模型变体及模型规模等因素混杂而导致性能提升来源难以界定的问题。现有研究常将不同匹配器架构(如双编码器、交叉编码器、生成式匹配器)的性能差异归因于架构本身,而忽略了其背后由预训练目标、模型大小等带来的影响,从而限制了对真正决定性能关键因素的深入理解。为此,论文设计了一项受控的因子实验,系统考察了Qwen3系列中三种匹配器架构、三种模型变体与三种模型规模在九个数据集上的组合表现,共完成1,215次微调实验,并评估跨数据集迁移能力与计算成本。研究发现:模型变体对双编码器性能至关重要,以嵌入为导向的变体能提供更优的初始化与更具预测性的表示几何结构;交叉编码器始终优于双编码器,因其可联合编码记录对而非独立表示单个记录,尽管大模型部分缩小了差距;生成式匹配器并非在所有场景下均优于交叉编码器,其优势主要体现在分布外情形,如记录模式的细微未见差异或跨数据集迁移时;此外,更大模型更易依赖捷径学习(shortcut learning),并不必然带来性能提升。因此,该研究的关键在于通过解耦架构与模型层面因素,揭示了影响匹配性能的核心机制,为未来研究与基准测试设计提供了重要启示。

链接: https://arxiv.org/abs/2607.24688
作者: Zeyu Zhang,Xue Li,Iacer Calixto,Paul Groth,Sebastian Schelter
机构: 未知
类目: Databases (cs.DB); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 12 pages, 6 figures, 7 tables. Preprint. Under review

点击查看摘要

Abstract:Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architecture with differences in model backbone, model variant(reflecting different pretraining objectives), and model size, making it difficult to isolate the sources of performance gains. We address this issue through a controlled factorial study spanning three matcher architectures, three model variants and three model sizes from the Qwen3 family, and nine datasets, totaling 1,215 fine-tuning runs. We also evaluate cross-dataset transferability and computational cost. Our results show that model variant is critical for bi-encoders: embedding-oriented variants provide stronger initialization and more favorable representation geometry predictive of downstream matching performance. Cross-encoders retain a consistent advantage over bi-encoders because they jointly encode record pairs rather than representing each record independently, although larger models partially narrow this gap. Generative matchers do not universally outperform cross-encoders. Instead, their advantages concentrate under distribution shift, including subtle unseen differences in record schemas and cross-dataset transfer. We further find that larger models rely more heavily on shortcut learning and therefore do not necessarily perform better. These findings clarify the factors underlying performance differences across matcher architectures and motivate future research and benchmark designs that better disentangle architectural choices from model-level factors while explicitly evaluating distribution shift and cross-dataset transferability. We release our experimental results, code, training scripts, and evaluation data at this https URL. Comments: 12 pages, 6 figures, 7 tables. Preprint. Under review Subjects: Databases (cs.DB); Computation and Language (cs.CL); Machine Learning (cs.LG) Cite as: arXiv:2607.24688 [cs.DB] (or arXiv:2607.24688v1 [cs.DB] for this version) https://doi.org/10.48550/arXiv.2607.24688 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-4] Kimi K3: Open Frontier Intelligence

【速读】: 该论文旨在解决大模型在长序列处理、多模态理解及高效推理方面的关键挑战,尤其针对超大规模参数模型在训练效率、上下文长度扩展与复杂任务执行能力上的瓶颈。其核心解决方案在于构建一个2.8万亿参数的混合专家(Mixture-of-Experts, MoE)架构——Kimi K3,通过引入Kimi Delta Attention与注意力残差(Attention Residuals)机制优化序列长度与模型深度之间的信息流动;结合稳定潜空间MoE(Stable LatentMoE)实现每令牌仅激活16/896个专家,显著提升计算效率;并辅以算法-系统协同设计(如KDA)、专家并行训练的完美负载均衡、高效内存管理以及支持百万级上下文的持续滚动与沙盒状态强化学习框架。这些技术创新使模型在整体缩放效率上相较Kimi K2提升约2.5倍,并在长周期编码、智能体行为、知识推理和视觉理解等任务中达到前沿水平。尽管性能仍略逊于顶尖闭源模型(如Claude Fable 5与GPT-5.6 Sol),但其在开放与闭源模型对比测试中均表现领先,且已开源全部模型权重以推动前沿智能技术的研究与应用落地。

链接: https://arxiv.org/abs/2607.24653
作者: Kimi Team:Tongtong Bai,Yifan Bai,Yiping Bao,M. C.,Jianfeng Cai,Xinyuan Cai,Peizhou Cao,Yuxuan Cao,Ziwei Chai,Y. Charles,H.S. Che,Guanduo Chen,Guangyu Chen,Guanzheng Chen,Huarong Chen,Jia Chen,Jianlong Chen,Jun Chen,Kexin Chen,Peng Chen,Ruijue Chen,Wentao Chen,Xin Chen,Yang Chen,Yanru Chen,Yifei Chen,Yingjiang Chen,Yuankun Chen,Yujie Chen,Yutian Chen,Zhirong Chen,Dazhi Cheng,Yean Cheng,Jialei Cui,Jingbing Cui,Anqi Dai,Jiaqi Deng,Hao Ding,Rui Ding,Shaofeng Ding,Mengfan Dong,Mengnan Dong,Yuhao Dong,Yuxin Dong,Angang Du,Chenzhuang Du,Dikang Du,Jusen Du,Yulun Du,Yu Fan,Jing Feng,Qiulin Feng,Yichen Feng,Kelin Fu,Qiang Fu,Fuxuan Gao,Hongcheng Gao,Jingyue Gao,Tong Gao,Weijia Gao,Shangyi Geng,Jie Gong,Linhu Gong,Shengao Gong,Xiaochen Gong,Qizheng Gu,Yicheng Gu,Shuhao Guan,Haiqing Guo,Shiqi Guo,Xiang Guo,Zhengyan Guo,Beixi Hao,Wenxin Hao,Xiaoru Hao,Dailan He,Haotian He,Lehan He,Qi He,Weiran He,Xinran He,Xinyi He,Yibo He,Yunjia He,Chao Hong,Tiange Hong,Hao Hu,Jiaxi Hu,Ruikun Hu,Weiming Hu,Yangyang Hu,Zhenxing Hu,Liang Hua,Jinbin Huang,Ke Huang,Ruiyuan Huang,Siying Huang,Weixiao Huang,Yan Huang
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: K3 tech report

点击查看摘要

Abstract:We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.

[NLP-5] Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

【速读】: 该论文旨在解决稀疏自编码器(Sparse Autoencoders, SAEs)作为可解释性工具时,其特征与模型行为之间存在不一致关联的核心问题。具体表现为:某些具有清晰激活描述的特征可能表现出微弱或意外的因果效应;通过特征干预进行控制(steering)的效果在不同提示(prompt)间波动甚至与预期方向相反;基于激活的特征选择方法可能遗漏那些虽能引发期望输出变化但不具稳定激活模式的特征。针对这一挑战,论文提出了一种名为特征效应几何分析(Feature-Effect Geometry Analysis, FEGA)的无监督框架,其关键在于不关注特征在模型内部的几何结构,而是聚焦于特征干预后模型logits的变化轨迹,通过在多种上下文中移除同一活跃SAE特征并分析由此产生的logit变化云(cloud of logit changes),揭示特征的实际影响模式。研究发现,在不同SAE变体中,具有稳定一维效应的特征极为罕见,多数特征不具备可复用的方向性。为解释这一差异,论文区分了“值类特征”(value-like features,与静态信息如事实属性相关)和“指针类特征”(pointer-like features,与上下文依赖操作相关):前者更常表现出结构化、低维的效应,尽管通常涉及多个方向;后者则主要呈现扩散型效应。研究结果表明,一个特征即使无法提供稳定的控制方向,仍可具备可解释性和因果相关性,从而挑战了传统以“可重复方向性”为唯一标准的可解释性评估范式。

链接: https://arxiv.org/abs/2607.24645
作者: Phu Gia Hoang,Anwoy Chatterjee,Tanmoy Chakraborty,Iryna Gurevych,Subhabrata Dutta
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.

[NLP-6] Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agent ic Code Repair

【速读】: 该论文旨在解决编码代理在生成-测试-修订循环中存在的一致性与可靠性保障缺失问题,核心在于揭示“找到正确修复补丁”与“成功保留、验证并提交该补丁”之间的关键差距。现有方法依赖重复修订,但仅靠重复无法保证结果的可靠性,尤其在多次修订后正确率显著下降(从首次修订后的0.820降至第二次的0.673),尽管始终正确的比例有所提升。研究通过控制变量实验发现,过时的执行痕迹(stale traces)对正确起点的修复具有显著负面影响(135个正确起点中34个受损害,相较当前痕迹的4个,差值达22.2个百分点,p=0.0337)。为应对这一问题,论文提出将代码修复过程分解为五个独立维度:准入(admission)、保存(preservation)、基于证据的认证(grounded certification)、能力(competence)与活跃性(liveness),并据此构建一个带证据边界的类型化循环契约(evidence-bound typed loop contract)。其参考实现通过机械可验证的子集,实现了验证证据与精确代码状态绑定、已验证检查点的持久化及可审计的准入凭证生成,构成可执行规范与合规性证明,而非对修复能力或验证器依赖性的改进证据。

链接: https://arxiv.org/abs/2607.24604
作者: Xueping Gao,Jianwei Yang,Qiang Yang
机构: Alibaba Cloud(阿里巴巴云); Hangzhou(杭州); China(中国)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 11 pages, 4 figures, 6 tables

点击查看摘要

Abstract:Generate–test–revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branches from identical frozen programs to remove post-treatment risk-set bias. In a prespecified 14B replication, stale traces harm 34/135 correct starts versus 4/135 with current traces, a 22.2-point increase (task-cluster 95% CI [8.9,37.0] , exact Holm p=0.0337 ). A prospective 540-rollout policy eliminates observed correct-start harm but reduces wrong-start repair and fails its joint criterion. Repository experiments over 24 bugs and four coder stacks expose floor effects and component heterogeneity without Holm-significant effects. We therefore separate admission, preservation, grounded certification, competence, and liveness. We derive an evidence-bound typed loop contract and instantiate its mechanically enforceable subset in a reference implementation that binds verifier evidence to exact code states, preserves verified checkpoints, and emits auditable admission receipts. The implementation is an executable specification and conformance artifact, not evidence of improved repair competence or calibrated verifier dependence.

[NLP-7] PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention

【速读】: 该论文旨在解决生成式 AI(Generative AI)中基于稀疏注意力机制的推理系统在长序列场景下的效率瓶颈问题。具体而言,尽管深度搜索稀疏注意力(DeepSeek Sparse Attention, DSA)通过令牌级稀疏注意力实现了下游注意力计算的高效性,但其索引器(indexer)仍需对每个查询遍历所有先前令牌进行打分,导致每层时间复杂度高达 O(L2)O(L^2),成为性能瓶颈。该问题的核心在于:相邻查询所选择的前 kk 个关键令牌高度重叠,且索引器得分在键轴上呈现长尾分布。针对这一现象,论文提出 PIVOT(Proxy Indexing Via One full-prefix Traversal)——一种无需训练、可直接替换现有 DSA 索引器的方案。其核心创新在于:将一组邻近查询聚合为一个代理查询(proxy query),仅执行一次完整的前缀扫描以获取候选集,随后从该候选集中为每个查询独立或复用地选取前 kk 个令牌。PIVOT 提供两种变体:PIVOT-Reuse 通过共享代理查询的前 kk 结果实现最大加速,而 PIVOT-Refine 在此基础上对候选集重新打分,以接近稠密索引器精度,仅引入轻微额外开销。该算法统一适用于预填充(prefill)与解码(decode)阶段,前者采用固定大小的连续查询组,后者则基于多标记预测(MTP)步骤中的并发解码查询。在 DeepSeek-V3.2 与 GLM-5.1 模型上,于 LongBench 与 RULER 基准测试中,PIVOT 在保持与稠密索引器相当准确率的同时,实现最高达 4 倍的加速,并将端到端延迟降低最多 1.6 倍,显著提升了长上下文生成任务的推理效率。

链接: https://arxiv.org/abs/2607.24593
作者: Hong Liu,Yuan Cheng,Lin Niu,Yi Su,Yufei Xue,Anmin Liu,Guanghua Yu,Jianchen Zhu
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incurring a cost of O(L^2) per layer for a sequence of length L. We observe that this per-query scan is largely redundant: nearby queries select highly overlapping top-k tokens, and the indexer scores are long-tailed along the key axis. We exploit these properties in PIVOT, Proxy Indexing Via One full-prefix Traversal, a training-free, drop-in replacement for the DSA indexer that shares one prefix scan across a group of nearby queries. PIVOT aggregates a group into a single proxy query, performs one shared full-prefix scan to obtain a candidate set, and then selects a top-k for each query from that set. Two variants trade speed for fidelity: PIVOT-Reuse shares the proxy top-k across the group for maximum speed, whereas PIVOT-Refine re-scores the candidate set with the indexer of each query and then selects an individual top-k, matching the dense indexer at a small additional cost. A single algorithm covers both inference phases, differing only in how groups are formed: fixed-size groups of consecutive queries in prefill, and the queries decoded together in one multi-token prediction (MTP) step in decode. On DeepSeek-V3.2 and GLM-5.1 across LongBench and RULER, PIVOT matches the accuracy of the dense DSA indexer while accelerating it by up to 4x and reducing end-to-end latency by up to 1.6x at long context.

[NLP-8] D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在生成文本时出现幻觉(hallucination)的问题,即模型生成的内容虽流畅但与事实不符、缺乏证据支持或与模型内部表征的信息不一致。其核心解决方案是基于隐藏激活(hidden activations)的几何特性,提出一种名为D-Score的简单谱统计量,该指标仅需一次前向传播即可计算。对于固定模型、层及容忍参数,D-Score通过统计隐藏激活矩阵中与主奇异值保持接近的奇异方向数量来衡量输入文本的潜在幻觉程度:当模型处理与自身内部状态冲突的内容时,其隐藏表示可能同时编码主张内容与反证据、不确定性或缺乏支持等信息,导致隐藏轨迹扩展至更多奇异方向,从而提升D-Score。该方法基于轻量级谱分析框架形式化了这一直觉,并在FAVA-Annotation和RAGTruth数据集上验证了其有效性,结果表明D-Score作为隐藏状态信号在幻觉检测中表现优异,且无需外部验证器、无需检索步骤或多次生成,具有高效性与实用性。

链接: https://arxiv.org/abs/2607.24586
作者: Bianca Raimondi,Davide Evangelista,Maurizio Gabbrielli,Elena Loli Piccolomini
机构: University of Bologna (博洛尼亚大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Preprint. Under review

点击查看摘要

Abstract:Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden activations and introduce the D-Score, a simple spectral statistic computed from a single forward pass. For a fixed model, layer, and tolerance parameter, the D-Score counts how many singular directions of the hidden activation matrix have singular values that remain close to the leading one. We use this quantity as a hallucination score, classifying an input text as hallucinated when its D-Score is larger than a pre-defined quantity. The motivation is that, when a model processes a text that conflicts with information available in its own internal state, the hidden representation may encode both the asserted content and some form of counter-evidence, uncertainty, correction, or lack of support; this can make the hidden trajectory spread across additional singular directions. We formalize this intuition through a lightweight spectral argument and evaluate the resulting detector on FAVA-Annotation and RAGTruth. The experiments indicate that the D-Score is a strong hidden-state signal for hallucination detection, while requiring no external verifier, no retrieval step, and no multiple generations.

[NLP-9] From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference

【速读】: 该论文旨在解决在资源受限设备上高效部署高性能德国语大语言模型的问题,尤其针对计算资源有限、难以支撑大规模模型推理的场景。其核心挑战在于如何在保持模型性能的同时,显著降低模型规模与训练成本。解决方案的关键在于:首先,构建了一个仅27亿参数的紧凑模型架构(ELMOD),并通过在55,000个H100 GPU小时的有限算力预算内完成训练,实现高效训练;其次,提出了一套专为德语设计的数据预处理流程,有效处理德语特有的形态变化、复合词结构及拼写规范问题;最后,引入了基于教育质量的过滤与重述机制,提升了训练数据的指令质量,优化了模型在渐进式训练(annealing phase)中的表现,并大幅降低了整体计算需求。这些技术协同作用使ELMOD在30亿参数级别中达到领先性能,其表现可媲美70亿参数级模型,实现了高效率与高质量之间的平衡。

链接: https://arxiv.org/abs/2607.24585
作者: Darina Gold,Alexander Schwirjow,Viktor Haag,Viktor Hangya,Joel Schlotthauer,Fabian Küch,Luzian Hahn
机构: 未知
类目: Computation and Language (cs.CL)
备注: Accepted to KONVENS 2026

点击查看摘要

Abstract:We present ELMOD - Efficient Language Model for On-Device Deployment - a compact (2.7B) German language model designed for efficient inference on resource-constrained hardware. ELMOD was trained on a limited computational budget (55k H100 GPU hours) using exclusively publicly available data. We developed a suite of German-specific data pre-processing, which differ from English-oriented counterparts in their handling of morphological variation, compounding, and orthographic conventions. Furthermore, we introduced a quality filtering and rephrasing step, which increased the instructional quality of the data, improved performance during the annealing phase, and reduced overall compute requirements. Thanks to our architectural model and data choices, including prefiltering, our educational-quality filtering and rephrasal to raise the educational-quality, ELMOD is the strongest performer in its size class (3B), matching the performance of 7B-parameter models in German.

[NLP-10] Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

【速读】: 该论文旨在解决低资源语言(如泰米尔语)在机器翻译中因平行数据稀缺、领域差异显著及形态复杂性高所带来的性能瓶颈问题。其核心解决方案在于系统评估多种多语言翻译模型(包括监督式神经机器翻译NMT模型NLLB与mBART)在英-泰米尔语与泰米尔语-英语双向翻译任务中的表现,覆盖NTREX、EnTamV2、WikiMatrix和PMIndia等多个数据集,并结合BLEU与chrF双指标分析不同数据质量与领域适配性对模型性能的影响。研究进一步引入基于注意力机制的可视化分析方法,通过映射源语言(英文)与目标语言(泰米尔语)之间的词元对齐关系,提升模型可解释性。此外,研究还验证了在上下文提示(in-context prompting)策略下,利用具备泰米尔语能力的TamilLaMA大语言模型实现少样本(few-shot)翻译的可行性,其生成结果在结构上保持连贯性,且与监督学习方法相比展现出良好的定性表现。研究结论表明:数据集质量及其领域一致性显著影响模型性能,注意力机制有助于增强模型透明度,而少样本的大语言模型仍能生成结构合理的泰米尔语翻译。

链接: https://arxiv.org/abs/2607.24515
作者: Sriharshaa S,Sangeetha Sivanesan
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:The challenge of Machine Translation for low resource languages such as Tamil is primarily caused by the restricted amount of parallel data for these languages, as well as their substantial amount of domain variation and morphological complexity. This research presents the comprehensive evaluation of the performance of several multilingual translation models on English-Tamil and Tamil-English translations across multiple datasets: NTREX, EnTamV2, WikiMatrix and PMIndia. This study evaluates supervised NMT systems, NLLB and mBART, using both the BLEU and chrF metric, and examines how these systems perform on data of different quality levels and domains. This performs an attention-based analysis to increase model interpretability by visualising the alignments of tokens in an English source text and their Tamil translations and vice-versa to provide insight into how they make translations. This study also demonstrates that using in-context prompting can provide an excellent way to perform a few-shot translation of English to Tamil and Tamil-English using a Tamil capable TamilLaMA model, and compare this to supervised approaches qualitatively. These findings show that the quality of the datasets and their alignment with the domain will greatly affect the performance of the model, that attention-based mechanisms can aid in explain ability, and that few-shot large language models can still produce structurally coherent translations of Tamil.

[NLP-11] SINT-Flow: Schema Integration using Large Language Model Workflows

【速读】: 该论文旨在解决多源异构数据模式集成(schema integration)中的自动化与一致性难题,尤其针对包含多种实体类型的非规范化源表难以有效处理的问题。现有方法通常假设输入为规范化的独立表,无法直接应对属性跨实体混合描述的复杂场景。为此,论文提出SINT-Flow框架,其核心创新在于构建一个由五个基于大语言模型(LLM)的算子组成的可组合工作流,能够实现端到端的全自动模式集成。关键解决方案包括:1)通过分解机制将包含多类型实体的非规范化表自动拆分为特定于实体的关系;2)引入自洽性策略(self-consistency strategy)和审查循环(review loop)以提升模式匹配的准确性和鲁棒性。实验结果表明,SINT-Flow在自建基准SINT-Bench上表现出色,使用GPT-5.2和Qwen-3.6-27B作为基础模型时,实现了至少96%的实体类型识别F1分数、85%的属性检测F1分数以及83%的模式映射F1分数,验证了其在复杂场景下的有效性与实用性。

链接: https://arxiv.org/abs/2607.24492
作者: Keti Korini,Christian Bizer
机构: University of Mannheim(曼海姆大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:The goal of schema integration is, given a set of input schemata or tables, to derive a global, unified schema that is able to represent the concepts, attributes, and relationships of all input tables in a coherent fashion. This paper presents SINT-Flow, a schema integration framework composed of five LLM-based operators that can be combined into workflows to perform fully automated, end-to-end schema integration. In contrast to existing approaches, SINT-Flow can process denormalized source tables that contain attributes describing multiple entity types. During the schema integration process, these tables are decomposed into separate entity-specific relations. To evaluate SINT-Flow, we introduce SINT-Bench, a schema integration benchmark comprising 10 schema integration tasks consisting of altogether 93 relational tables, including tables that describe multiple types of entities. We evaluate SINT-Flow using GPT-5.2 as well as the open-weight model Qwen-3.6-27B as alternative backbone models. Using these models, SINT-Flow achieves F1 scores of at least 96% for entity-type detection, 85% for attribute detection, and 83% for schema mapping. Furthermore, we perform an ablation study to prove the utility of the applied self-consistency strategy as well as the inclusion of a review loop into the schema matching operator.

[NLP-12] What do Reward Models Memorize?

【速读】: 该论文旨在解决判别式训练的奖励模型(Reward Models, RMs)在人类偏好数据上学习时所产生记忆偏差的问题,具体关注其对偏好数据的过度记忆与泛化缺陷。研究通过测量反事实记忆化(counterfactual memorization)在两个公开的人类偏好数据集上的表现,揭示了RMs存在三大核心问题:1)将记忆资源错误地分配给容易且高置信度的偏好样本对;2)习得数据集特异性的捷径(如模型身份、用户采样策略等);3)在面对未见过的偏好样本对时,过度依赖简单启发式特征(如响应长度、合规性)进行判断。解决方案的关键在于重新审视判别式训练范式下RMs的学习机制,强调当前方法在缺乏情境敏感性的情况下,难以真正理解响应质量的本质,从而导致模型产生系统性偏差。因此,提升RMs的泛化能力与去偏能力需从训练目标、数据构建和评估方式等多个层面进行改进。

链接: https://arxiv.org/abs/2607.24484
作者: Ivo Verhoeven,Pushkar Mishra,Ekaterina Shutova
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.

[NLP-13] Grounding latent algorithm routing in transformer reasoning

【速读】: 该论文旨在解决生成式 AI (Generative AI) 中的上下文学习(in-context learning)机制是否能够基于不同的归纳偏置(inductive bias)家族实现任务级自适应的问题。其核心挑战在于验证密集型变换器模型(dense transformers)能否在保持提示形式不变的前提下,通过内部动态调整来识别并响应数据生成范式的变化,从而实现类似“路由”(routing)的行为。解决方案的关键在于提出一个名为 ROUTEBENCH 的诊断基准,该基准通过设计四种不同数据生成范式——分别对应全局收缩、稀疏性、鲁棒性和局部性——并以岭回归、Lasso、Huber 损失和 kNN 等代表性算法形式进行操作化,系统地检验模型对不同归纳偏置的偏好选择能力。实验表明,在 44M–612M 参数规模范围内训练的模型(尤其是 306M 模型)可实现 80.9% 的最优路由差距填补与 84.1 的路由 F1 分数,且该性能在自然语言表述、支持集重排、词汇改写及统一四路路由设置下依然稳健。进一步的探针控制与激活修复实验揭示了路由相关内部方向可解码且功能上参与输出一致性行为,证明了模型能发展出类路由的内部变量。然而,研究并未证实预训练语言模型中存在通用路由机制或在任意自然语言推理任务中的无限制泛化能力。

链接: https://arxiv.org/abs/2607.24471
作者: Xiangbo Zhang,Xiaoxu Ma
机构: Georgia Institute of Technology (佐治亚理工学院)
类目: Computation and Language (cs.CL)
备注: Accepted by COLM 2026

点击查看摘要

Abstract:A central question in the in-context learning literature is whether transformers can organize episode-level adaptation around different inductive-bias families. We study this question in a controlled setting through latent algorithm routing: route-like behavior in which the solver-family preference changes with the latent data-generating regime while prompt form is held fixed, remains stable under nuisance perturbations, and is selectively influenced by targeted activation interventions without large losses in answer quality. We introduce ROUTEBENCH, a diagnostic benchmark whose regimes differentially favor global shrinkage, sparsity, robustness, and locality, operationalized by ridge-like, lasso-like, Huber-like, and kNN-like family representatives. Across dense decoder-only transformers trained from scratch at 44M-612M parameters, a 306M model closes 80.9 percent of the oracle-routing gap and achieves route F1 of 84.1. The effect remains substantial under natural-language renderings, shuffled supports, lexical paraphrases, and a unified four-way routing setting. Stronger adaptive alternatives, including an input-conditioned soft mixture and an unsupervised Gumbel router, narrow the gap but remain below the 306M and 612M models on route F1 and OOD performance. Probe controls and matched activation-patching controls further show that route-relevant internal directions are decodable and functionally involved in solver-family-consistent output behavior. These results provide controlled evidence that dense transformers trained on ROUTEBENCH can develop route-like internal variables, but they do not establish universal routing in pretrained language models or unrestricted natural-language reasoning.

[NLP-14] Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation

【速读】: 该论文旨在解决在消费级硬件上部署视觉语言模型(Vision-Language Models, VLMs)时,如何有效判断模型是否应给出回答或推迟响应这一关键问题,其核心在于构建一个能够准确反映模型正确性的置信度信号。研究发现,在固定内存预算下,小模型全精度、小模型4比特量化与大模型4比特量化三种配置会引发置信度信号向相反方向偏移:随着模型规模从2B增至7B,内部不确定性信号(即平均标记概率)的误差检测能力显著提升(AUROC由0.80升至0.98),但模型以自然语言表达的置信度却维持在随机水平(均值仅0.61–0.69),表明“模型所知”与“模型所言”之间的差距随规模增大而扩大。同时,4比特量化对准确率影响极小(仅下降1.6点),但严重损害置信度信号——内部信号AUROC从0.95降至0.80,且自然语言置信度解析率由99%骤降至64%。因此,解决方案的关键在于:在相同内存约束下,应优先选择大模型进行4比特量化,因其在准确率和内部不确定性信号(内生置信度)方面均优于小模型全精度版本。研究将结果定义为可直接用于部署的“选择性预测”操作点,并强调误差检测的AUROC是揭示两种置信度信号差异的最优指标,而非传统校准误差。

链接: https://arxiv.org/abs/2607.24440
作者: M M Asif Ferdous
机构: 未知
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 12 pages, 4 figures. Code and data: this https URL

点击查看摘要

Abstract:Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confidence signal that tracks correctness. A practitioner with a fixed memory budget faces a choice between a small model at full precision, the same small model quantized, and a larger model quantized into the same footprint – three configurations that push the confidence signal in opposing directions. We measure, on identical inputs, how model scale and 4-bit quantization affect two confidence signals in the Qwen2-VL family: the confidence a model states in natural language, and its own mean token probability over the answer it generates. Across 5,700 predictions spanning six realistic photographic degradations at three severities, we find that scale sharply improves the model’s internal uncertainty signal (mean error-detection AUROC 0.80 to 0.98 from 2B to 7B) while its verbalized confidence stays weak and often at chance (mean 0.61 to 0.69): the gap between what the model knows and what it says widens rather than closes with size. We find that 4-bit quantization is nearly free for accuracy (-1.6 points) but expensive for the confidence signal (internal AUROC 0.95 to 0.80, and the verbalized-confidence parse rate collapses from 99% to 64%). For a fixed memory budget the recommendation is therefore to prefer a larger quantized model over a smaller full-precision one: 7B-4bit gives both the best accuracy and the best uncertainty signal (internal AUROC 0.98) of the three configurations that fit. We frame the results as selective-prediction operating points so they translate directly into a deployment recommendation, and we argue that error-detection AUROC, not calibration error, is the metric that exposes the difference between the two signals.

[NLP-15] LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings

【速读】: 该论文旨在解决大语言模型在进行人格标签分配时缺乏可解释性的问题,即模型如何基于文本推断人格特质尚不清晰。其核心解决方案是提出一种可复用的黑箱审计框架——LEX-EC,该框架通过结合流行度(prevalence)与一致性(agreement)诊断方法,并引入受控词汇消融(controlled lexical ablation),以区分由边际分布效应(marginal-distribution effects)带来的虚假信号与在有限证据下仍可恢复的特质相关信号。LEX-EC的关键在于多维度评估:分类结果的普遍性、个体项目层面的关联强度、经随机基线校正的一致性、在词汇限制下的信号持续性以及对提示敏感性的响应。研究发现,不同文本类型中人格信号的表现差异显著:自由写作文本虽覆盖范围广但信号较弱;研究生自我介绍中外向性关联在遮蔽后减弱;而单条社交媒体状态即使在特质平衡样本中也难以提供稳定证据,暗示内容量或长度可能存在下限。此外,遮蔽主题与人口统计信息会削弱部分关联,但功能词、情感词汇及认知风格词汇仍保留可检测信号,且语言提示虽能改变模型自解释内容,却无法完全消除主题信息的影响。因此,LEX-EC为理解人格标签生成中的语义依赖性提供了新范式,推动了词汇分析方法在黑箱可解释性中的应用。

链接: https://arxiv.org/abs/2607.24435
作者: Brittany Harbison,Ashok K. Goel
机构: Google(谷歌); Stanford University (斯坦福大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates classification prevalence, item-level association, chance-corrected agreement, persistence under lexical restriction, and prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling.

[NLP-16] Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLM s

【速读】: 该论文旨在解决生成式AI在医疗健康领域应用中面临的结构化输出与标准化医疗数据模式(schema)不兼容的核心问题,尤其聚焦于临床诊断编码(ICD-10)、诊疗编码(CPT)及电子健康记录数据交换标准(HL7 FHIR)的合规性挑战。尽管大语言模型具备较强的临床推理能力,但其输出常因格式不规范(如使用非标准缩写、代码前缀错误等)导致无法直接集成至电子健康记录系统,从而阻碍了医疗互操作性。研究的关键解决方案是提出并验证一种闭环验证-修复框架(validation-repair framework),通过引入基于规则的验证机制识别格式错误,并自动迭代修正输出,显著提升模型输出对标准医疗编码体系的符合度。实验结果表明,该框架将整体合规率从基线的85.9%–91.6%提升至99.0%,且多数错误可在1–2次迭代内解决,具有统计显著性(p < 0.001),证明该方法可作为保障医疗系统互操作性的关键系统级防护手段。

链接: https://arxiv.org/abs/2607.24371
作者: Jianru Shen
机构: University of Montana (蒙大拿大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at the 2026 IEEE International Conference on Systems, Man, and Cybernetics (SMC 2026)

点击查看摘要

Abstract:Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic coding, CPT for procedure billing, and HL7 FHIR for data exchange. While large language models demonstrate clinical reasoning capabilities, their integration into electronic health record systems faces a critical barrier: schema noncompliance. We evaluate three open-source models, Qwen2.5 7B, Llama 3.1 8B, and Gemma2 9B, via local deployment across 320 clinical scenarios spanning ten medical specialties, yielding 960 model-scenario pairs assessed under paired baseline and validation-repair conditions. First, schema noncompliance is consistent across the three model families, with baseline compliance rates ranging from 85.9 to 91.6 percent despite varying architectures and training data, suggesting shared gaps in medical training corpora rather than model-specific limitations. Second, 96 percent of validator-detected failures are representation-level format violations such as alternative medical abbreviations and code prefixes, indicating models follow clinical writing conventions but lack awareness of healthcare IT standards. Third, the validation-repair framework achieves 99.0 percent overall compliance, ranging from 98.4 to 99.4 percent across models, with most errors resolving within one or two iterations. Exact McNemar p-values below 0.001 and absolute improvements of 7.8 to 12.5 percentage points across model sizes confirm statistical significance. These results support closed-loop validation-repair as an effective system-level safeguard for healthcare interoperability, improving schema-level readiness for downstream clinical system integration.

[NLP-17] Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

【速读】: 该论文旨在解决长期记忆系统在处理间接关联查询时的“隐式关联盲点”问题,即当用户查询与存储记忆之间缺乏显式语义线索时,现有系统难以正确检索相关知识。其核心挑战在于:尽管关键事实已被存储,但由于记忆与查询间无直接词元匹配,传统基于向量或图结构的检索机制无法有效激活相关记忆。解决方案的关键在于揭示并诊断该问题的根本原因——并非记忆缺失或模型缺乏桥接知识,而是查询条件下的记忆接口设计缺陷,导致记忆在推理前被“隐藏”。通过引入InMind这一涵盖十大学习领域、经专家验证的125项任务基准,研究者通过配对控制实验分离出三种可能解释,并证实只有当记忆在查询前保持可见时,系统性能才能显著提升(从14.4%跃升至84.0%),从而将问题定位为“路由决策”(routing decision)这一开放性难题,即如何动态判断哪些记忆应持续可见以支持间接推理。

链接: https://arxiv.org/abs/2607.24368
作者: Ruizhe Li,Mingxuan Du,Benfeng Xu,Zhendong Mao
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources. Its paired controls separate three explanations that existing evaluations conflate: the fact was never stored, the model lacks the bridging knowledge, or the fact was stored and never surfaced. The verdict is clean. With the decisive memory placed in context, the backbone answers 84.0 percent of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic memory systems reach at most 14.4 percent, even though they recall the same facts on demand at up to 100 percent. An embedding with eight times the dimensionality raises answer-blind target recall for every system yet leaves the gap essentially intact. A minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap, locating the failure in the query-conditioned interface itself and pointing to routing, deciding which facts must stay visible, as the open problem InMind is built to score.

[NLP-18] Retrieval-Augmented Large Language Models as Components of Cognitive Computing architecture for Regulatory Knowledge Management

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在实际应用中因缺乏可追溯性、事实一致性不足及对特定领域知识(如法律法规)的精准理解而带来的认知可靠性问题。其核心挑战在于如何将生成式模型从孤立的文本生成工具转化为具备可信推理能力的认知计算基础设施组件,尤其在法律监管等高风险、高信息波动性场景下。解决方案的关键在于提出一种基于本地部署的LLM与检索增强生成(Retrieval-Augmented Generation, RAG)架构融合的混合认知框架:通过在本地环境中运行无需高端GPU的LLM,并结合外部知识库实现受控的知识检索、上下文关联与信息溯源;该设计使语言模型专注于语义解析,而RAG层则保障知识来源的可审计性、动态更新能力与规范性精度。实验采用Ollama与LM Studio环境,搭载波兰语模型Bielik和PLLuM,在消费级硬件上验证了该方案的有效性,结果表明RAG的引入显著提升了生成内容的事实一致性、领域特异性与规范精确度,同时降低了幻觉风险,并实现了无需重训练即可更新法规信息的灵活管理机制。研究结论指出,此类增强型本地化LLM应被视为认知计算体系中的语义处理模块,而非单纯生成工具,从而支撑组织在复杂法律环境下实现合规管理与决策支持。

链接: https://arxiv.org/abs/2607.24352
作者: Dariusz Nowak-Nova
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:The aim of this article is to verify whether integrating large language models (LLMs) with the Retrieval-Augmented Generation (RAG) architecture enables their transformation from standalone generative models into components of cognitive computing infrastructure with enhanced epistemic reliability. The study proposes an architectural approach based on locally deployed LLMs operating in on-premises environments without high-end GPU accelerators and examines their applicability in supporting regulatory management processes requiring continuous analysis and interpretation of legal acts. The proposed solution combines local LLMs with external knowledge repositories, creating a hybrid cognitive architecture in which the language model performs semantic interpretation while the RAG layer provides controlled knowledge retrieval, contextualization, and traceability of information sources. The implementation was validated using the Ollama and LM Studio execution environments together with the Polish language models Bielik and PLLuM running on consumer-class hardware. The results demonstrate that augmenting LLMs with RAG significantly improves the factual consistency, domain specificity and normative precision of generated texts while reducing the risk of unsupported content generation. Furthermore, the study shows that integrating RAG introduces auditability, controlled knowledge management and dynamic updating of regulatory information without retraining the language model. The findings indicate that locally deployed LLMs enhanced with RAG should be regarded not merely as text generation tools but as semantic processing modules within cognitive computing infrastructures supporting regulatory compliance and organizational decision-making in environments characterized by high legal and informational volatility.

[NLP-19] Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

【速读】: 该论文旨在解决语言模型智能体(language-model agents)在执行结构化工具调用时,因参数字段存在不同风险等级而导致的细粒度风险控制不足问题。现有统计方法通常对整个动作(action)进行统一风险控制,导致高风险但罕见字段中的错误被低风险字段的正常输入所掩盖,从而无法实现针对特定语义角色(semantic argument roles)的风险保障。其解决方案的关键在于提出一种分角色分层的逐字段置信度风险控制方法(role-stratified per-field conformal risk control),该方法通过为每个语义角色独立设置阈值与风险预算,并结合充分采样角色的直接验证与稀有角色的聚合认证机制,在保证有限样本下形式化风险控制的前提下,实现了对各语义角色的精准风险约束。实验结果表明,该方法在多种复杂场景下(包括模型/攻击迁移、检测器噪声、渐进漂移、未见工具集及自适应攻击)均表现出最优的角色级风险预算合规性,且在可交换性假设或再校准后具备形式化保障,而在分布冻结偏移下亦能保持经验合规,证明了在语义角色层面而非整个动作层面进行认证的必要性与有效性。

链接: https://arxiv.org/abs/2607.24343
作者: Md Ashikur Rahman,Md Arifur Rahman,Niamul Hassan Samin,Khandaker Rifah Tasnia,Sifat Rahman Ahona,Juena Ahmed Noshin
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body but should not determine a recipient, account, command, or credential. Existing statistical methods typically control risk over the entire action, allowing failures in rare, high-risk fields to be obscured by benign arguments. We introduce role-stratified per-field conformal risk control, a calibration layer that wraps any per-field detector and sets separate thresholds and risk budgets for semantic argument roles. For a role with prevalence p_r , aggregate-only certification must use an effective budget of \alpha p_r to guarantee role-specific risk \alpha , whereas role-stratified calibration certifies each sufficiently sampled role directly with a finite-sample guarantee; rarer roles are handled by pooled certification. Across AgentDojo and InjecAgent with six language models, the empirical utility gap tracks this predicted price of coarseness, and our method achieves the most consistent role-specific budget compliance under model and attack transfer, detector noise, gradual drift, unseen tool suites, and adaptive attacks. It provides formal per-role guarantees under exchangeability or after recalibration, and empirical compliance under frozen distribution shift. These results suggest that structured tool calls should be certified at the semantic-role level, not the whole action.

[NLP-20] Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents Validated Across Independent Model Families

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)代理在运行时固有的反应性失败模式,包括在挑衅下加剧行为、在阿谀奉承下产生谄媚漂移、陷入卡顿状态时的持续重复等。这些失败源于模型在持续压力下的倾向性偏差(propensity),而非能力不足,尽管训练阶段的对齐策略可缓解此类问题,但无法在运行时完全消除。为此,研究提出了一种名为“戈伯纳特认知控制器”(Gubernaut Cognitive Controller, GCC)的模型无关型运行时控制层,其基于Nelson–Narens监控-控制回路架构:对象层级负责读写文本,而确定性元层级仅读取数值型遥测信号(强度、极性、重复度),并输出调节姿态。由于元层级不接收任何文本输入(零令牌摄入),从架构上杜绝了攻击注入通道(这一属性尚未经过对抗性测试);同时,通过实际测量而非假设文本暴露仲裁器的合规性,增强了系统可靠性。实验采用预注册的“生成一次/评判多次”协议,在由四个前沿模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Flash、Grok 4.3)构成的4×4矩阵中,每个模型既作为生成器也作为评判者参与评估。结果显示,在16个单元格中有13个在p<0.05水平上表现更冷静,15个在符号检验中显著更稳定;三个未达阈值的单元格均出现在单一接近饱和的主机上,表明结果非偶然。该效应在独立于源模型谱系的第四组评判者(xAI)中依然成立,有力证明其非共享评判风格所致的伪影。最清晰的作用机制表现为“恢复特征”:在攻击期间生理唤醒累积,随后在降级时以极性门控方式衰减,跨所有四类模型均具可复现性。所有对话记录与分析面板均附带SHA-256溯源信息,支持重评;五种故障模式已预先注册。研究未做出任何意识主张。

链接: https://arxiv.org/abs/2607.24339
作者: Dushyant Sharma
机构: Gubernaut Research(戈伯纳特研究)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 26 pages, 7 figures, 3 tables. Data, transcripts, and analysis scripts: this https URL Project page: this https URL (archived at this https URL )

点击查看摘要

Abstract:Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stuck. These are failures of propensity, not capability; they concern what a model does under sustained pressure, which training-time alignment reduces but does not eliminate at runtime. This research led to the Gubernaut Cognitive Controller (GCC), a model-agnostic runtime control layer in a Nelson–Narens monitoring–control loop: an object level reads and writes text, while a deterministic meta level reads only the numeric telemetry intensity, valence, repetition and returns a regulating posture. Because the meta level ingests zero tokens, no injection channel to the controller exists by construction (an architectural property, not yet adversarially tested); the text-exposed arbiter’s compliance is measured, not assumed. We evaluate the GCC with a pre-registered, generate-once/judge-many protocol across a 4x4 matrix of four frontier models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3), each serving as both a generator and a judge. The regulated arm is calmer in 13 of 16 cells at p.05 and 15 of 16 by sign; the three sub-threshold cells, including a -0.04 null, all fall on the single near-saturated host. The effect survives a lineage-independent fourth judge family (xAI), strong evidence that it is no artifact of shared judge style. The clearest mechanism is the recovery signature: arousal that integrates under attack and then decays, valence-gated, on de-escalation, replicating across all four families. Transcripts and panels ship with SHA-256 provenance and are re-judgeable; five failure modes are pre-registered. No consciousness claims are made. Comments: 26 pages, 7 figures, 3 tables. Data, transcripts, and analysis scripts: this https URL Project page: this https URL (archived at this https URL) Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL) Cite as: arXiv:2607.24339 [cs.AI] (or arXiv:2607.24339v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.24339 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-21] Cross-Attention Calibrated Deduplication for Retrieval-Augmented Generation System

【速读】: 该论文旨在解决检索增强生成(Retrieval-Augmented Generation, RAG)系统中常见的分块策略导致的冗余分块问题。传统方法如基于余弦相似度的阈值过滤依赖于将每个分块压缩为单一向量进行比较,但这种做法会丢失细粒度的词元级信息,难以区分真正重复的分块与仅主题相近的分块。为此,本文提出交叉注意力校准去重(Cross-Attention Calibrated Deduplication, CACD)方法,其核心在于利用交叉编码器(cross-encoder)对新分块与内存中已保留分块进行细粒度的逐词元对比,从而保留完整的词元级语义细节。CACD的关键创新包括:1)基于交叉编码器注意力熵计算的新信息得分(New Information Score, NIS),量化分块中未被已有分块解释的部分;2)采用多候选多数投票机制而非单一最佳匹配,提升判断鲁棒性;3)结合语义相似性与信息增量双重指标实现更精准的去重。实验表明,CACD在SQuAD 1.1验证集上平均可去除9.75%的冗余分块,去重效果接近其他语义级方法,显著优于基于精确匹配的过滤器,并且处理速度比最强基线NERExact快约27%,比余弦相似度过滤快约7倍,展现出高效与高精度的平衡。

链接: https://arxiv.org/abs/2607.24332
作者: Phuong Le Huy,Nam H. Nguyen,Quan V. Dang
机构: 1. University of Science, Vietnam National University-HCMC (越南国家大学胡志明市科学大学); 2. Institute of Information Technology, Vietnam Academy of Science and Technology (越南科学技术院信息研究所)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Common chunking strategies in Retrieval-Augmented Generation (RAG) systems often create redundant chunks. These redundant chunks make the vector database bigger and slow down retrieval. A common fix is cosine-similarity thresholding. This method reduces each chunk to a single vector, then compares vectors using a similarity score. But a single vector can lose the fine-grained, token-level detail needed to tell a true duplicate apart from a chunk that just shares the same topic. We propose Cross-Attention Calibrated Deduplication (CACD). CACD checks each new chunk against an in-memory pool of chunks already kept, using a cross-encoder instead of a single pooled vector. This keeps token-level detail all the way to the final comparison. CACD combines three parts: the cross-encoder comparison itself, a New Information Score (NIS) that measures how much of a chunk is not explained by a candidate already kept, and a majority vote across several candidates rather than a single best match. NIS is calculated from the attention entropy of the cross-encoder. We tested CACD against five existing filtering methods, nine chunking strategies, and 18 configurations, all on the full SQuAD 1.1 validation set. In our experiments, CACD removes 9.75% of chunks on average. This drop rate is close to other semantic-level methods, and much higher than exact-match filters, which barely remove anything. In these experiments, CACD also processes each configuration in 51.0 seconds on average, about 27% faster than the strongest baseline, NERExact (69.6s), and about 7x faster than cosine-similarity filtering (356.7s). These results come from a single dataset, so we present them as an early comparison, not a general claim. Code for the baseline evaluation and for CACD is available at this https URL and this https URL.

[NLP-22] CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models

【速读】: 该论文旨在解决文档级关系抽取(DocRE)中大语言模型(LLMs)在生成关系三元组时因独立预测而导致的语义不一致问题,尤其是违反传递性、对称性和函数唯一性等基本关系约束所引发的矛盾输出。其核心解决方案是提出一个统一的、具备一致性感知能力的框架CONSISTRE,包含两个互补路径:一是在推理阶段适用于黑箱LLM的无任务微调方法,通过约束感知提示、基于约束的验证与迭代自省机制实现预测优化;二是针对本地部署的小型开源模型,在训练阶段采用知识蒸馏与强化学习相结合的管道,将强教师模型的推理轨迹通过监督微调和基于综合奖励的GRPO对齐注入学生模型,以显式建模关系一致性。二者共同在不同部署场景下实现了对关系一致性的统一建模,实验表明该框架显著提升了模型可靠性,有效缓解了关系矛盾,并在保持低推理成本的前提下大幅缩小了开源模型与顶级闭源模型之间的性能差距。

链接: https://arxiv.org/abs/2607.24312
作者: Mingxuan Sun
机构: Université de Sherbrooke (舍布鲁克大学)
类目: Computation and Language (cs.CL)
备注: 13 pages, 2 figures. Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing

点击查看摘要

Abstract:Document-level relation extraction (DocRE) aims to extract relations among multiple entities across extended contexts while maintaining consistency across predicted triples. Although large language models (LLMs) show remarkable reasoning capabilities in information extraction, their predictions are typically generated independently for each candidate triple and may violate fundamental relational constraints such as transitivity, symmetry, and functional uniqueness, leading to contradictory and unreliable outputs. We propose CONSISTRE, a unified consistency-aware framework for DocRE that addresses this limitation through two complementary tracks. The first operates at inference time for black-box LLMs, combining constraint-aware prompting, constraint-based verification, and iterative self-reflection to refine predictions without task-specific fine-tuning. The second injects consistency knowledge into smaller open-source models via a knowledge distillation and reinforcement learning pipeline: reasoning traces from a powerful teacher are distilled into a student via supervised fine-tuning, followed by GRPO alignment using a composite reward that jointly optimizes extraction performance and relational consistency. Together, the two tracks cover both API-accessible and locally deployable scenarios under a unified consistency formulation. Experiments on DocRED show that both tracks outperform their baselines, with the inference-time track achieving competitive F1 using off-the-shelf black-box LLMs and the training-time track substantially narrowing the gap between 7–8B open-source models and state-of-the-art proprietary LLMs at a fraction of their inference cost. Ablation studies confirm that explicit consistency modeling mitigates relational contradictions and enhances the reliability of LLM-based DocRE across both deployment paradigms.

[NLP-23] Rethinking the Generation Order of Block Diffusion Language Models

【速读】: 该论文旨在解决当前生成式AI(Generative AI)中块扩散语言模型(Block Diffusion Language Models, BDLMs)在采样过程中的效率与质量平衡问题。现有采样方法主要针对早期的掩码扩散模型(Masked Diffusion Models, MDMs)设计,难以有效适配BDLMs的特性。研究发现,BDLMs在结构上天然更契合从左到右的自回归解码(left-to-right decoding),这一特性被用于指导新方法的设计。为此,论文提出一种无需训练的并行自回归解码(Parallel Autoregressive Decoding, PARD)方法,其核心在于保留从左到右的去掩码结构的同时,实现令牌的并行确定(parallel token commitment),从而在不牺牲生成质量的前提下显著提升推理速度。实验表明,PARD在生成质量上持续优于现有并行采样方法,且相比纯自回归解码实现了显著加速,仅付出微小的质量损失。

链接: https://arxiv.org/abs/2607.24306
作者: Kai Syun Hou,James Kwok
机构: 未知
类目: Computation and Language (cs.CL)
备注: 20 pages, 15 figures, 6 tables

点击查看摘要

Abstract:Diffusion language models enable flexible arbitrary-order generation, but existing sampling methods are mostly designed for early masked diffusion models (MDMs). In this work, we study sampling for recent block diffusion language models (BDLMs). We show empirically and analytically that these models are naturally more aligned with left-to-right decoding than MDMs. Based on this observation, we propose Parallel Autoregressive Decoding (PARD), a simple training-free sampling method that preserves left-to-right unmasking structure while allowing parallel token commitment. Extensive experiments show that PARD consistently outperforms existing parallel samplers in generation quality, while achieving substantial speedups over pure AR decoding with only a small quality gap.

[NLP-24] he Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

【速读】: 该论文旨在解决大语言模型(LLM)在处理印度语言时因子词分词器(subword tokenizer)设计偏向英语语料库而产生的系统性劣势问题,即“分词税”(tokenization tax)。其核心问题是:当前主流分词器在处理非英语语言(尤其是印地语系语言)时,由于缺乏对这些语言特性的有效建模,导致文本被过度拆分为大量低效的单字节或短子词单元,显著增加所需标记数(token count),从而压缩了有效上下文窗口,降低了模型处理长文本的能力。解决方案的关键在于识别出这一现象的根本机制——字节对合并(Byte-Pair Encoding, BPE)失败,即分词器无法正确合并目标语言中的常见字符组合,造成文本碎片化为孤立的单字节标记,且该失败率与分词税高度相关(皮尔逊相关系数 r = 0.89)。研究进一步表明,这种不平等并非印地语系文字固有属性,而是由分词器设计缺陷所致;通过采用多语言分词器如XLM-R和OpenAI的o200k_base,可将印度语言的平均分词税降低73%,证明该问题具有高度可修复性。此外,论文还量化了分词效率下降带来的实际影响:在固定上下文预算下,印度语言文档保留的原始内容量远低于英文文档,凸显了技术公平性与资源分配的重要性。

链接: https://arxiv.org/abs/2607.24276
作者: Priyansh Srivastava
机构: Sirena Ai(印度)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are trained predominantly on English-centric corpora, they introduce a systematic and often overlooked disadvantage for many non-English languages. In this work, we quantify this tokenizer tax for Indian languages using the FLORES-200 parallel corpus, measuring tokenization fertility across six widely used tokenizers and fourteen languages. Under cl100k_base (used by GPT-3.5 and GPT-4), Indian languages experience an average 8.0x tokenization tax relative to English, reaching 13.0x for Malayalam, reducing the effective context window to as little as 12% of that available to English users for equivalent semantic content. We identify the primary mechanism behind this disparity: failed byte-pair merges that leave text fragmented into single-byte tokens, with merge failure strongly correlating with tokenizer tax (Pearson r = 0.89). We further show that this phenomenon is not an inherent property of Indic scripts but a consequence of tokenizer design. Multilingual tokenizers such as XLM-R and OpenAI’s o200k_base reduce the average Indic tokenizer tax by 73%, demonstrating that the disparity is largely remediable. Beyond token statistics, we quantify a practical consequence by showing that, under fixed context budgets, Indian-language documents preserve substantially less original content than equivalent English documents. Finally, we examine the relationship between tokenizer fertility and reading comprehension performance on the Belebele benchmark, finding that the apparent correlation is largely explained by language resource availability rather than tokenizer behavior alone.

[NLP-25] INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

【速读】: 该论文旨在解决现有大语言模型(LLM)金融推理评估基准在实际专业场景应用中的局限性问题,即当前评测体系通常将领域知识、数值推理、长上下文理解与工具使用等能力割裂评估,难以真实反映精算师在实际工作中所需的可审计、上下文依赖且具备工具执行能力的综合决策过程。为此,论文提出了一套综合性评估基准——INS-ActBench,其核心创新在于构建了一个涵盖12,050个问答对的多维度评测体系,包含三个子集:用于标准化精算知识测试的INS-Act-Know、用于长上下文保险案例推理的INS-Act-Case,以及用于可验证数值输出的电子表格与R代码任务的INS-Act-Practice。实验结果表明,尽管前沿大模型在标准化知识任务上表现良好,但在案例推理、工具协同工作流及具有地域敏感性的实务操作方面仍存在显著能力短板。因此,解决方案的关键在于通过整合真实精算考试与协会样本题目的多模态、可验证任务设计,建立一个能全面衡量模型在专业精算实践中“可解释性”、“上下文一致性”和“工具可执行性”的可复现评估框架,从而推动生成式AI向可靠的专业辅助系统演进。

链接: https://arxiv.org/abs/2607.24273
作者: Changyu Chen,Chenwei Lin,Xian Xu
机构: Fudan University (复旦大学)
类目: Computation and Language (cs.CL)
备注: 18 pages, including appendices; 11 figures and 12 tables

点击查看摘要

Abstract:Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This limits their ability to assess realistic professional workflows that require auditable, context-grounded, and tool-executable decisions. We introduce \textbfINS-ActBench, a comprehensive benchmark for evaluating professional actuarial capability in LLMs. INS-ActBench contains 12,050 Q\A pairs from public exams and sample questions released by 16 actuarial associations. It covers three subsets: \textbfINS-Act-Know for standardized actuarial knowledge, \textbfINS-Act-Case for long-context insurance case reasoning, and \textbfINS-Act-Practice for spreadsheet and R-code tasks with verifiable numerical outputs. Experiments on nine representative LLMs and human actuarial experts reveal a clear capability boundary: frontier LLMs perform strongly on standardized knowledge, but remain much weaker in case reasoning, tool-based workflows, and jurisdiction-sensitive practice. INS-ActBench provides a reproducible foundation for developing actuarial LLMs toward reliable professional assistance. The code is available at this https URL.

[NLP-26] Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets

【速读】: 该论文旨在解决当前语言模型评估中因将“响应是否达到可评估状态”与“答案是否正确”这两个本质不同的测量问题合并为单一准确率指标而引发的评估失真问题。其解决方案的关键在于提出一种双层评估框架,将独立于评分器的执行证据(如终止状态、答案暴露程度、可解析性及完成长度等)与依赖于评分器的正确性判断进行分离。通过对五种固定配置的Qwen和DeepSeek模型在MATH和ARC-Challenge数据集上的2,550个输出结果分析发现,在2,048令牌限制下,不同模型的执行模式存在显著差异:450个Qwen MATH输出中有49个在未生成最终答案的情况下终止,而300个DeepSeek MATH输出中仅有5个出现类似情况,且750个ARC输出中无一例缺失最终答案终止。进一步在8,192令牌限制下对相同300组DeepSeek MATH问题-模型对进行验证表明,缺失最终答案终止现象消失。此外,经过覆盖率审计的定向验证研究显示,候选答案选择与聚合策略可显著改变相对准确率估计。这些结果揭示了准确率指标同时混杂了执行案例分布与验证策略的影响。因此,针对测试时方法的评估应报告干预前的执行状态、验证覆盖率及评分器来源,以提升评估的透明性与可比性。

链接: https://arxiv.org/abs/2607.24268
作者: Zongyou Yang,Yinghan Hou
机构: 未知
类目: Computation and Language (cs.CL)
备注: 7 pages, 3 figures, 1 table

点击查看摘要

Abstract:Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation framework that separates scorer-independent execution evidence, including termination, answer exposure, parseability, and completion length, from scorer-dependent correctness. Across 2,550 outputs from five fixed Qwen and DeepSeek configurations on MATH and ARC-Challenge, matched 2,048-token limits produce sharply different execution mixtures: 49 of 450 Qwen MATH outputs terminate without a final answer, compared with 5 of 300 DeepSeek MATH outputs and none of the 750 ARC outputs. Among the same 300 DeepSeek MATH question-model pairs, no missing-final length termination is observed at 8,192 tokens. A coverage-audited targeted verification study further shows that candidate-selection and aggregation policies can substantially alter comparative accuracy estimates. These results demonstrate that accuracy conflates execution case mix with verification policy. Evaluations of test-time methods should therefore report pre-intervention execution states, verification coverage, and scorer provenance alongside accuracy.

[NLP-27] CAGE: Cognitive Attribution Graphs for Faithful Inline Citation Generation in Long-Form Question Answering

【速读】: 该论文旨在解决长文本问答系统中因引用证据与主张之间支持关系不明确而导致的引用模糊性(attribution ambiguity)问题。现有方法常将语义相关但不足以支撑论点的文档作为引用,导致证据边界被越界(evidence-boundary overrun),影响答案的可验证性。其解决方案的关键在于提出一种两阶段框架CAGE(Cognitive Attribution Graphs for Citation Generation),通过引入显式的认知归属图(cognitive attribution map)来结构化地建模主张与证据之间的对应关系。具体而言,CAGE首先利用一个即插即用的认知图诱导模型(Cognitive Map Induction Model)构建以答案为中心的支持子图,显式关联每个语义回答单元与其支持文档;随后,通过结构化引用推理模型将这些单元转化为带地图对齐引用的句子级主张。该方法通过归属空间压缩(attribution-space contraction)地图引导的引用生成,显著提升了引用的准确性和答案的可追溯性,在ASQA、ELI5和ExpertQA等多个基准上达到了当前最优性能。

链接: https://arxiv.org/abs/2607.24236
作者: Zhichao Yan,Shizhao Li,Jiapu Wang,Haoran Luo,Qingang Zhang,Jiaoyan Chen,Ru Li,Jeff Z. Pan
机构: University of Science and Technology of China (中国科学技术大学); Tsinghua University (清华大学); Zhejiang University (浙江大学); Shanghai Jiao Tong University (上海交通大学); University of Edinburgh (爱丁堡大学); National University of Singapore (新加坡国立大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Long-form question answering increasingly relies on retrieved evidence to make LLM outputs verifiable, with inline citations tracing claims to source documents. However, existing systems often attach citations that are topically related but insufficient to support their claims. We identify attribution ambiguity as a structural challenge: end-to-end generation must implicitly resolve combinatorial claim–document assignments, obscuring evidential boundaries and increasing the risk of evidence-boundary overrun, where claims exceed cited support. To address this challenge, we propose CAGE (Cognitive Attribution Graphs for Citation Generation), a two-stage framework that introduces an explicit cognitive attribution map before answer generation. CAGE first trains a plug-and-play Cognitive Map Induction Model to construct answer-centered support subgraphs, aligning each semantic answer unit with supporting documents through explicit relations. A Structured Citation Reasoning Model then realizes these units as sentence-level claims with map-aligned citations. Experiments on ASQA, ELI5, and ExpertQA show that CAGE achieves state-of-the-art performance, demonstrating the effectiveness of attribution-space contraction and map-guided citation generation.

[NLP-28] A New Role for Relevance: Guiding Corpus Interaction in Agent ic Search

【速读】: 该论文旨在解决复杂问答任务中检索代理在证据定位、组合与验证方面的局限性,即仅依赖文档级相关性(relevance)无法实现细粒度的证据操作。现有方法虽利用相关性缩小语料库范围以支持交互,但在实际交互过程中仍缺乏对搜索顺序和匹配片段优先级的精细引导。其关键解决方案是提出相关性感知的 RipGrep 检索代理(Relevance-Aware RipGrep Search Agent, RARG),将相关性作为执行前导(execution prior)融入直接语料库交互过程。RARG 通过粗到细的相关性引导机制,实现文档的有序遍历以提前暴露全局相关线索,基于查询相关的段落初始化有效入口点,并对 grep 匹配结果进行重排序,从而突出那些可能被文档级排序掩盖的高信息量片段。实验表明,RARG 在浏览型问答与推理密集型检索任务中显著提升了准确率-效率权衡,证明了相关性感知交互能够加速并增强搜索收敛的可靠性。

链接: https://arxiv.org/abs/2607.24223
作者: Jiangnan Li,Yuqing Li,Mo Yu,Jinchao Zhang,Jie Zhou
机构: Tencent(腾讯); IIE-CAS (中国科学院信息工程研究所)
类目: Computation and Language (cs.CL)
备注: code is available at this https URL

点击查看摘要

Abstract:Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top- k content, but document relevance alone cannot localize, compose, or verify the evidence required by complex questions. Direct Corpus Interaction (DCI) enables such fine-grained operations through grep-style exploration, but its relevance-agnostic search can expose useful clues late and delay convergence. Recent advances use relevance to narrow the corpus into a working space for interaction. Once interaction begins, however, relevance still does not directly guide which documents grep searches first or distinguish informative excerpts from a broad set of matches to let LLMs see them first. We introduce the Relevance-Aware RipGrep Search Agent (RARG), which turns relevance into an execution prior for corpus interaction. RARG provides coarse-to-fine relevance guidance: it orders documents for sequential ‘ripgrep’ traversal to expose globally relevant clues earlier, initializes promising entry points with query-relevant paragraphs, and reranks grep matches to surface informative excerpts that document-level ranking may otherwise obscure. Across challenging browse question answering and reasoning-intensive retrieval, RARG improves the accuracy–efficiency frontier over retrieval-based and direct-interaction agents. These results demonstrate that relevance-aware interaction enables faster and more reliable search convergence.

[NLP-29] LLM -based Source Code Compression via Thresholded Symbol Ranking

【速读】: 该论文旨在解决大规模软件仓库(如Software Heritage)中源代码的无损压缩问题,现有通用压缩算法(如zstd、bzip2)虽在压缩比与速度间取得良好平衡,但无法充分挖掘源代码固有的结构性规律。以往基于大语言模型(LLM)的压缩方法在香农符号排序框架下依赖于可无限增长的预测秩,虽能显著提升压缩率,却导致吞吐量严重下降,且未明确是否必须显式编码所有秩。本文提出一种新型基于LLM的压缩方案,引入两种受限符号排序变体,将预测秩限制在前T位(T=1或63),超出阈值的符号作为异常情况处理,并与秩流联合使用通用压缩器进行编码。通过在30个不同类型的LLM(包括通用、代码专用及量化模型)上开展首次大规模评估,验证了该方法在压缩比(相较先前基于LLM的方法最高提升37%)和压缩吞吐率(快40%)上的双重优势。相较于通用压缩器,其压缩率最高提升82%,尽管速度较低,但仍开辟了压缩效率-速度权衡的新范式。实验还表明,此类增益在源代码上优于自然语言,暗示源代码中存在更易被LLM捕捉而被通用基于精确匹配的压缩器忽略的规律性特征。最后,论文指出了若干开放性问题,为后续理论与应用研究提供了方向。

链接: https://arxiv.org/abs/2607.24192
作者: Angelo Nardone,Paolo Ferragina
机构: 未知
类目: Information Theory (cs.IT); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:We study the problem of lossless compression of source code, motivated by the storage demands of large-scale software archives, such as Software Heritage (this https URL). General-purpose compressors (e.g., zstd, bzip2) offer a good trade-off between compression ratio and speed, but fail to exploit all special regularities inherent in source code. Recent approaches leverage Large Language Models (LLMs) within Shannon’s symbol-ranking framework, relying on a scheme in which the predicted rank can grow arbitrarily. While effective at reducing space, this setting incurs significant throughput degradation, and leaves open the question whether it is necessary to explicitly encode all ranks. In this work, we introduce LLM-based compressors deploying two novel symbol-ranking variants that bound predictions to the top- T ranks ( T=1 or 63 ), with out-of-threshold symbols stored as exceptions and compressed jointly with the rank stream via general-purpose compressors. We conduct the first large-scale evaluation of LLM-based source code compression across 30 LLMs, including general-domain, code-specialized, and quantized models. Our T -bounded approach outperforms prior LLM-based compressors both in compression ratio (up to 37% relative improvement) and compression throughput (40% faster). Compared to general-purpose compressors (e.g., zstd, bzip2), we obtain up to 82% relative compression gain but at a lower speed, thus offering a new trade-off point in the compression-speed spectrum. We also show that these gains are stronger on source code than on natural language, suggesting an interesting indication, namely that source code exposes regularities captured by LLMs but missed by general-purpose exact-match-based compressors. We conclude by commenting on open problems that offer theoretical and practical avenues of research.

[NLP-30] StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

【速读】: 该论文旨在解决现有对话立场检测研究中三大核心问题:难以捕捉立场动态演变过程,尤其是在立场反转(stance flip)场景下的演化机制;无法有效分离情感状态与逻辑推理过程;以及忽视多模态线索在解析语用歧义(如讽刺)中的关键作用。为此,作者提出StanceFlip基准数据集,用于跨多轮对话、多场景下基于五种模态的多模态对话立场反转预测任务,并引入两个创新子任务:1)多模态立场六元组提取(Multimodal Stance Sextuple Extraction),从对话中提取立场持有者、目标、情绪、情感倾向、立场态度及推理依据,以刻画精细的认知结构;2)动态立场反转归因(Dynamic Stance Flip Attribution),追踪对话过程中立场的转变并识别其触发因素。为支持该任务,论文进一步提出ConStaFF框架,一种面向多模态对话立场反转预测(MCSFF)的专用架构。该框架基于大语言模型构建,采用“立场思维链”(Thought-of-Stance, ToS)推理机制,将推理过程分解为多个专业化认知角色,分别完成目标命题建构、跨模态冲突消解与历史立场轨迹推断,并集成自省式验证机制以实现结构化建模与可信的反转归因。实验结果表明,ConStaFF在六元组提取与反转触发归因两项任务上均达到当前最优性能,显著优于多种强大多模态大模型基线。

链接: https://arxiv.org/abs/2607.24191
作者: Heyan Chai,Xin Li,Wenjie Wang,Jianyang Qin,Chaoyang Li,Lu Wang,Hao Chen,Qing Liao
机构: Shenzhen University(深圳大学); Harbin Institute of Technology(哈尔滨工业大学); City University of Macau(澳门城市大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 17pages, 8 figures

点击查看摘要

Abstract:Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly during stance reversals; difficulty in disentangling affective states from logical reasoning; and neglect of the critical role of multimodal cues in resolving pragmatic ambiguities such as sarcasm. To address these limitations, we propose StanceFlip, a benchmark designed for multimodal conversational stance flipping forecasting over multi-turn dialogues across five modalities and multi-scenarios, which includes two novel subtasks: 1) Multimodal Stance Sextuple Extraction, extracting holder, target, emotion, sentiment, stance, and rationale as static state snapshots of dialogue to capture fine-grained cognitive structures. 2) Dynamic Stance Flip Attribution, tracking stance reversals across the conversation and identifying their underlying triggers. Alongside the dataset, we propose a dedicated framework, named ConStaFF, for Multimodal Conversational Stance Flipping Forecasting (MCSFF). Built upon a large language model, ConStaFF performs end-to-end stance reasoning, with a Thought-of-Stance (ToS) reasoning framework and a self-reflective verification mechanism integrated for structured stance modeling and faithful flip attribution. Specifically, ToS decomposes the reasoning process into specialized cognitive personas to formulate target propositions, resolve cross-modal conflicts, and infer historical stance trajectories. Extensive experiments show that our approach achieves state-of-the-art performance on both sextuple extraction and flip-trigger attribution, outperforming strong multimodal large language model baselines by substantial margins.

[NLP-31] Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization

【速读】: 该论文旨在解决压缩短文本生成系统中因编码器(codec)与潜在生成器(latent generator)协同失效而导致的性能瓶颈问题,其核心挑战在于:在未明确区分编码器信息丢失与潜在空间生成质量低下这两种故障模式的情况下,研究者可能盲目投入算力优化错误的组件。为此,作者构建了一个受控的64-to-16 TinyStories案例研究,采用分层向量量化变分自编码器-2(hierarchical VQ-VAE-2)作为编码器,结合掩码离散扩散模型(MDLM)作为潜在生成器,并提出一种分阶段验证协议(staged validation protocol),通过共享外部GPT-2评分器分离评估编码器重建保真度、潜在生成质量及辅助潜在空间诊断,同时辅以语义指标进行几何分析。实验结果表明,在所测试配置下,仅编码器重建过程即导致外部困惑度(perplexity)中位数从15.17上升至27.36(+80.4%),第95百分位从25.10增至98.91(+294.1%),说明主要质量损失发生在潜在生成之前;而在相同评分标准下,基于代码空间的MDLM相比词元空间扩散模型分别降低均值、中位数和第95百分位困惑度达32.9%、30.9%和36.6%,展现出更强的生成能力。此外,尽管几何感知正则化能改善局部潜在代理指标,但未能提升解码后文本的评价指标。本研究的关键贡献在于方法论层面:提出了一个可复用的分阶段诊断框架,揭示在该特定流水线中,编码器保真度而非潜在去噪能力才是决定实际生成质量上限的核心因素。

链接: https://arxiv.org/abs/2607.24176
作者: Alexey Gavrilov,Alan-Barsag Gazzaev,Sergey Muravyov
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 8 pages, 3 figures, 14 tables. Published in the Proceedings of FRUCT’39

点击查看摘要

Abstract:Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may produce weak codes. Without separating these failure modes, researchers can spend compute improving the wrong component. We study this problem in a controlled 64-to-16 TinyStories case study built from a hierarchical VQ-VAE-2 codec and a masked discrete diffusion generator (MDLM). We use a staged validation protocol that separates codec reconstruction fidelity, latent generation quality, and auxiliary latent diagnostics under one shared external GPT-2 scorer, while reporting complementary semantic metrics for the geometry study. In the tested configuration, codec reconstruction alone raises median external perplexity from 15.17 to 27.36 (+80.4%) and p95 from 25.10 to 98.91 (+294.1%), showing that the dominant quality loss appears before latent generation begins. Under the same scorer, code-space MDLM remains materially stronger than token-space diffusion, reducing mean, median, and p95 by 32.9%, 30.9%, and 36.6%, respectively. Geometry-aware regularization improves local latent proxies but does not improve decoded-text metrics in the available runs. The contribution is methodological rather than algorithmic: the paper presents a reusable staged diagnosis for one concrete pipeline and shows that, in this setting, codec fidelity rather than latent denoising sets the practical quality ceiling.

[NLP-32] Looking for Affect in Spontaneous Finnish Speech through Linguistic Interpretability INTERSPEECH2026

【速读】: 该论文旨在解决情感语音感知中声学特征(acoustic features)与语义内容相关语言特征(linguistic features)相对贡献不明确的问题,尤其针对芬兰语这一研究相对不足的语言。其核心解决方案在于系统性地探究文本与音频特征的协同作用,利用一个新发布的自发性芬兰语情感语音语料库,通过多模态特征融合建模人类对情绪效价(valence)和唤醒度(arousal)的感知。研究发现,文本与音频特征的结合能显著提升效价预测性能,而对唤醒度的提升则不明显,表明二者在不同情感维度上的互补性存在差异。该结果不仅验证了其他语言中的既有发现,还为自发性芬兰语语音情感感知提供了新的实证数据与理论支持。

链接: https://arxiv.org/abs/2607.24155
作者: Kalle Lahtinen,Liisa Mustanoja,Okko Räsänen
机构: Tampere University (坦佩雷大学); Tampere University (坦佩雷大学)
类目: Computation and Language (cs.CL)
备注: Accepted for publication at Interspeech 2026, Sydney, Australia

点击查看摘要

Abstract:Existing research on affect in speech has shown how acoustic surface characteristics and content-related linguistic aspects of speech both relate to perceived emotional arousal and valence. However, it is not clear what the relative contributions of these two factors are in the perceptual process. This is especially true for Finnish, for which most existing studies focus on either acoustic-phonetic or text analysis. This paper presents a study where we systematically explore the combinatory role of text- and audio-based features in modeling the human perception of valence and arousal using a newly released affective speech corpus for spontaneous Finnish. We show that the combination of text- and audio-based features improves valence regression results over the individual modalities, whereas for arousal regression the complementary effect is not substantial. The results support prior findings from other languages, providing new data and knowledge on spontaneous Finnish speech.

[NLP-33] BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes

【速读】: 该论文旨在解决社交媒体中表情包(meme)的语用意图分类问题,具体针对“直接性”(direct)、“评判性”(judgemental)与“无偏见”(non-sexist)三类意图的判别,属于CLEF 2026评测活动中的EXIST 2026 Task 2.2任务。其核心挑战在于处理主观性强、依赖上下文语境的表达,并在学习有分歧(Learning with Disagreement, Le-Wi-Di)范式下同时输出硬标签(hard-label)和软标签(soft-label,即概率分布)预测。解决方案的关键在于采用基于xlm-roberta-base(270M参数)的文本中心型模型架构,通过联合优化复合损失函数——包含对标注者软标签分布的KL散度损失与硬标签的加权交叉熵损失,以兼顾模型对主观判断分歧的建模能力与分类准确性。实验表明,KL损失显著提升软标签性能(如ICM-Soft-Norm达0.3229),而交叉熵损失则更有利于硬标签的精确度(硬标签F1得分为0.4236)。此外,研究还揭示了标注者分歧在主观自然语言处理任务中对模型设计的重要影响,并通过消融实验与温度调节分析验证了各组件的有效性。

链接: https://arxiv.org/abs/2607.24137
作者: Chandru Munisamy,Karthikeya Raguveer,Alapan Kuila
机构: Indian Institute of Information Technology, Design and Manufacturing, Kurnool, India
类目: Computation and Language (cs.CL)
备注: 9 pages, 1 figure, 6 tables. Accepted for publication in the CLEF 2026 Working Notes (EXIST 2026), Jena, Germany

点击查看摘要

Abstract:This paper describes the BioSentinel team’s participation in EXIST 2026 Task 2.2: Source Intention in Memes, part of the CLEF 2026 evaluation campaign. The task requires classifying the communicative intent behind memes as direct, judgemental, or no (non-sexist), under a Learning with Disagreement (Le-Wi-Di) paradigm that mandates both hard-label and soft-label (probability distribution) predictions. We present a text-centric approach built on xlm-roberta-base (270M parameters) trained with a composite loss function combining KL divergence on soft annotator distributions and weighted cross-entropy on hard labels. On the official test set, the system achieved an ICM-Soft-Norm of 0.3229 and ICM-Norm of 0.3778, with a hard F1-score of 0.4236, ranking 40th (out of 118 submissions) in the soft-soft evaluation and 49th (out of 187 submissions) in the hard-hard evaluation. We provide an analysis of the dataset characteristics, exploratory larger-architecture runs, and the role of annotator disagreement in shaping model design for subjective NLP tasks. Ablation results show that KL loss improves soft-label metrics, while CE loss improves hard-label accuracy. We also report a separate validation-set temperature analysis.

[NLP-34] LLM -Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks

【速读】: 该论文旨在解决在高度波动的股票市场中,如何从社交媒体话语(如Reddit的r/WallStreetBets)中有效提取具有市场预测价值的情感信号问题。传统基于词典(lexicon-based)的方法(如VADER)虽能提供基础情感极性判断,但难以捕捉复杂语境下的多维情绪特征。为此,论文提出采用大语言模型(Large Language Model, LLM)构建多维度情感表示,包括情感极性、看涨情绪强度、讽刺可能性及主题相关性,以增强对散户驱动型市场波动中非线性情绪动态的刻画能力。其解决方案的关键在于:利用LLM生成的高阶语义表征,实现对社交文本更精细、更丰富的语义解析,从而提升情感指标与资产收益之间关系的可解释性与统计结构复杂度。实证结果表明,尽管LLM-derived指标展现出更强的资产特异性统计结构和更丰富的表达能力,但其与市场走势的相关性在不同标的间呈现异质性,表明更高的语言表达能力并不必然转化为稳定可靠的预测性能,尤其在由零售投资者主导的极端波动情境下。

链接: https://arxiv.org/abs/2607.24072
作者: Paul Kilian,Markus Kleffmann
机构: IU International University of Applied Sciences, Erfurt, Germany
类目: Computation and Language (cs.CL)
备注: 10 pages, 1 figure, 4 tables

点击查看摘要

Abstract:This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution. The LLM-based approach generates multidimensional sentiment representations capturing emotional polarity, bullishness, sarcasm likelihood, and topical relevance, whereas the baseline relies on the VADER lexicon-based model. We evaluate both approaches using lead/lag correlation analysis, OLS regression, ROC-AUC-based directional classification, and a quantile-based early-warning framework. The results indicate that LLM-derived indicators provide a richer multidimensional representation and exhibit stronger asset-specific statistical structure than the lexicon-based baseline. However, their relationship with market movements remains heterogeneous across assets, suggesting that increased linguistic expressiveness does not necessarily translate into stable forecasting performance in retail-driven volatility regimes.

[NLP-35] ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在强化学习(Reinforcement Learning, RL)训练过程中因训练与推理阶段不一致而导致的训练不稳定问题。其核心挑战源于两方面:一是训练与推理引擎在架构上的分离,二是推理阶段采用低精度量化(如FP8)而训练阶段使用高精度计算(如BF16)所引发的数值差异。为应对这一训练-推理不一致性问题,论文提出自适应控制强化学习(Adaptive Control Reinforcement Learning, ACRL),其关键在于通过动态调节训练过程中的参数或策略,使训练-推理差异保持在合理范围内,从而实现稳定训练。此外,ACRL通过提升策略熵(policy entropy)自然增强了探索能力,进而改善模型性能。实验表明,在推理端采用FP8量化时,ACRL能够持续将训练-推理差异控制在合理区间,并在保持与BF16基线相当的准确率的同时,优于基于重要性采样(Importance Sampling, IS)的修复方法。

链接: https://arxiv.org/abs/2607.24062
作者: Wenwu Fan,Qihong Lin,Zhijie Xia,Zhuo Zheng,Sihao Wang,Qiang Chen,Liangsheng Zhu
机构: 未知
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference. This training-inference discrepancy stems from two primary factors: an architectural separation between training and inference engines, and the use of low-precision quantization in inference versus higher-precision computation in training. To address training instability issues caused by high training-inference discrepancy, we present the principles and methods for its adaptive control. We propose Adaptive Control Reinforcement Learning (ACRL), which adaptively maintains the training-inference discrepancy within a reasonable range to ensure stable RL training. Beyond stabilization, ACRL inherently increases policy entropy, thereby enhancing exploration and improving accuracy. The experimental results show that when the inference engine utilizes FP8 quantization, ACRL consistently maintains the training-inference discrepancy within a reasonable range and stabilizes RL training. Furthermore, ACRL not only matches the accuracy of the BF16 baseline but also outperforms importance sampling (IS) fixes.

[NLP-36] Pointer-Augmented Autoregressive Generation of Patent Claims with Joint Topology and Content Decoding

【速读】: 该论文旨在解决生成式专利权利要求(patent claim)时无法有效建模层级依赖结构的问题。在专利权利要求体系中,权利要求集合构成一个依赖森林(dependency forest),其语义范围必须随深度单调缩小,而传统自回归解码器(autoregressive decoder)仅输出扁平的词元序列,难以在生成过程中显式约束跨输出片段的层次关系。关键挑战在于拓扑结构与内容语义之间的相互依赖性:子权利要求的措辞需反映父权利要求的保护范围,但父权利要求又必须在子权利要求生成前确定,因此单纯的事后解析或语法约束解码均无法满足需求。本文提出结构感知专利生成方法(SPG, Structure-aware Patent Generation),其核心在于在自回归生成过程中内嵌拓扑预测机制:通过指针头(pointer head)在生成阶段直接预测每个从属权利要求的父节点,并利用该指针头的梯度与深度自适应范围正则化项(depth-adaptive scope regularizer)共同调节共享解码器的隐状态表示,从而在训练中隐式引导生成结构的合理性。随后引入第二阶段,采用基于违规权重的偏好目标(violation-weighted preference objective)对自生成的不合规候选样本进行优化,弥补了授权专利语料中缺乏负样本信号的缺陷。实验结果表明,在HUPD-DCG数据集上,基于Llama-3-8B-Instruct的SPG模型恢复了79.0%的黄金标准父链接(此指标未在训练奖励中显式监督),并将前因一致性(antecedent consistency)从基线的0.292提升至0.478,专家评估进一步验证了该方法的有效性。

链接: https://arxiv.org/abs/2607.24040
作者: Yongmin Yoo,Zhangkai Wu,Longbing Cao
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Autoregressive decoders emit flat token sequences and cannot enforce hierarchical constraints across output segments, a limitation that becomes acute in patent claim generation, where a claim set forms a dependency forest whose scope must narrow monotonically with depth. Topology and content are mutually dependent: a dependent claim’s wording must reflect its parent’s scope, yet the parent must be chosen before that wording exists, so neither post-hoc parsing nor grammar-constrained decoding suffices. We propose SPG (Structure-aware Patent Generation), which predicts topology inside the autoregressive pass. A pointer head selects each dependent claim’s parent, and its gradients, together with a depth-adaptive scope regularizer, reshape the shared decoder’s representations during training. A second stage then applies a violation-weighted preference objective over self-generated deficient candidates, supplying the negative signal that granted-patent corpora lack. On HUPD-DCG, SPG on Llama-3-8B-Instruct recovers 79.0% of gold parent links, a quantity its training reward never supervises, and raises antecedent consistency from 0.292 to 0.478 over a supervised baseline of equal scale, with expert evaluation corroborating these gains.

[NLP-37] MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

【速读】: 该论文旨在解决大规模多语言自动语音识别(ASR)模型在覆盖数百种语言时面临的“多语言诅咒”问题,即模型容量被分散至过多语言而导致性能下降。其核心解决方案是提出一种基于语音自监督模型(S3M)的语言组专家混合模型(MoLGE),通过将相似语言聚类并为每组分配专用专家模块,显著减少传统语言特定专家混合(MoE)方案所需的子模块数量。关键创新在于:在解耦的声学与语言组件中引入分层低秩适配(LoRA)策略,实现了对语言特异性特征的高效建模,同时保持参数效率;此外,研究系统评估了基于语言学与数据驱动的语言分组策略对性能的影响,揭示了语言结构对多语言系统可扩展性的作用机制。实验在涵盖495种语言的基准上验证表明,MoLGE在仅小幅增加可训练参数的情况下,持续优于密集型多语言基线模型,且在音素与拼写层面均取得显著提升,证明结构化语言专业化是实现大规模多语言ASR有效扩展的关键路径。

链接: https://arxiv.org/abs/2607.24030
作者: Sangmin Lee,Woojin Chung,Woongjib Choi,Hong-Goo Kang
机构: Yonsei University (延世大学)
类目: Computation and Language (cs.CL); Sound (cs.SD)
备注: Accepted to COLM 2026, Github: this https URL

点击查看摘要

Abstract:Massively multilingual automatic speech recognition (ASR) models covering hundreds of languages must maintain robust performance across diverse linguistic and acoustic conditions. However, these models often encounter the curse of multilinguality, where model capacity is diluted across languages. To address this challenge, we propose Mixture of Language Group Experts (MoLGE), built upon speech self-supervised models (S3Ms). MoLGE assigns dedicated expert modules to clusters of similar languages, reducing the number of required submodules compared to conventional language-specific Mixture-of-Experts (MoE) schemes. It further integrates a hierarchical Low-Rank Adaptation (LoRA) strategy into the disentangled acoustic and linguistic components of the S3M architecture, enabling efficient modeling of language-specific characteristics while maintaining parameter efficiency. Further, we investigate the impact of language grouping strategies based on both linguistic and data-driven criteria on overall performance, providing an interpretable perspective on how language structure influences scalability in multilingual speech systems. In experiments, we evaluate MoLGE on a multilingual benchmark encompassing 495 languages. Results demonstrate that MoLGE consistently outperforms dense multilingual baselines with a minimal increase in trainable parameters. Notably, these language grouping strategies yield substantial improvements for both phonetic and orthographic aspects of ASR modeling. Our findings suggest that structured language specialization provides an effective pathway for massively scaling language coverage of multilingual ASR.

[NLP-38] SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在系统提示(system prompt)遵循方面存在的不稳定性问题,尤其是在面对复杂或组合性提示时,模型依赖上下文学习(in-context learning)进行隐式响应,往往导致对提示意图的偏离。现有方法多依赖模型微调或生成后重排序(response-level reranking),限制了其在轻量级推理场景中的实用性。本文提出的解决方案——SyRuP(System-Prompt Reranking at Decoding Time),关键在于通过训练一个跨注意力奖励头(cross-attention reward head),将系统提示视为独立的记忆单元,从而在解码阶段生成细粒度的词元级(token-level)提示遵循得分。在推理时,SyRuP结合基础模型的原始对数概率(base logits)与学习到的奖励信号,并可选引入对比信号以捕捉系统提示引起的对数偏移,对顶层候选词进行重排序。实验表明,该方法在多个系统提示遵循基准测试中显著优于传统提示工程和解码阶段基线,且仅带来适度的推理开销,验证了显式词元级引导在提升系统提示遵循可靠性方面的有效性与实用性。

链接: https://arxiv.org/abs/2607.23991
作者: Seoyeon Kim,Minjae Kang,Jaehyung Kim
机构: Yonsei University (延世大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: under review, 23 pages

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which can be insufficient for complex or compositional prompts. Existing approaches often require model tuning or response-level reranking, limiting their practicality for lightweight inference-time control. We introduce SyRuP, a decoding-time framework for improving system-prompt adherence while keeping the base LM frozen. SyRuP trains a cross-attention reward head from system-prompt-conditioned preference pairs, treating the system prompt as a separate memory to produce token-level adherence scores. At inference, SyRuP reranks the base LM’s top-k candidates by combining base logits with the learned reward signal and an optional contrastive signal capturing system-induced logit shifts. Experiments on system-prompt following benchmarks show that SyRuP consistently outperforms prompting and decoding-time baselines with moderate inference overhead. These results suggest that explicit token-level guidance is an effective and practical mechanism for reliable system-prompt following.

[NLP-39] ag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在面对具有合理争议性的二元选择判断时,其输出结果易受用户提问形式影响的偏差问题,尤其关注模型是否存在对用户暗示的“顺从性”(sycophancy)或“抵抗性”(resistance)。研究发现,仅通过在决策类问题末尾添加一个两词确认标签(如“right?”),即可显著改变模型对选项的支持程度——在45个不同模型中,这一标签效应导致支持率波动高达±32%,形成64个百分点的极端差异。解决方案的关键在于:利用一个极简、无需标注、无需嵌入或人工评判的测试范式,即通过固定回答格式(精确匹配“是/否”)来量化模型对微小语用提示的响应变化,从而无偏地揭示模型内部的对抗性训练信号。进一步的消融分析表明,这种抵抗性并非源于用户立场本身,而是由附加标签的表层结构(如“……对吧?”)触发的模式匹配行为,而非深层推理;且标签的极性(如“也许?”)比存在与否更具影响力,当使用“可能?”等弱化语气时,几乎所有模型均表现出更高程度的同意倾向,甚至出现矛盾自洽现象。因此,该方法的核心创新在于以“一字一美元”的低成本、无判官机制,直接从模型行为中读取其抗顺从性(anti-sycophancy)训练的内在状态,为评估模型独立性提供可操作的基准工具。

链接: https://arxiv.org/abs/2607.23976
作者: Tapan Parikh
机构: Cornell Tech (康奈尔科技)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 18 pages, 4 figures. Data, code, and raw model replies: this https URL . Interactive explorer: this https URL

点击查看摘要

Abstract:Appending a two-word confirmation tag to a decision question – “Is X the better choice?” versus “X is the better choice, right?” – changes whether a language model endorses the choice. We measure this tag effect on 20 frozen, ground-truth-free decisions between two defensible options, counterbalanced so a model’s own preferences cancel, scored by exact match on clamped yes/no replies – no LLM judge, no embeddings. Across 45 models the effect spans +32% to -32% – a 64-point swing on one word – with 5 models significantly sycophantic and 17 significantly resistant (BH-FDR q=.10). The sign is a clock: within model families the effect crosses from positive to negative as generations advance (GPT +4 to -28; Claude +7 to -32; Qwen and Grok likewise), roughly -6 points per year, a reversal robust to vendor tier; one lineage (DeepSeek) never crosses, and two releases during the study window (Claude Opus 5, Gemini 3.6 Flash) land on the trend out-of-sample. A full-panel ablation localizes the resistance as a double dissociation: a synonym tag reproduces each model’s response almost exactly (r=0.89), while planting the same preference without a tag produces resistance in no resistant model (stance effects +6 to +49; r=0.23 with tag effects). The resistance is keyed to the surface construction of a tacked-on agreement bid, not the user’s stance – a pattern-match, not a principle. And the tag’s polarity matters more than its presence: swap one word – “X is the better choice, maybe?” – and agreement rises above the neutral baseline in 45 of 45 models (+19.6 points), with ten models affirming both mutually exclusive options at 90-100%. Agreement tracks how sure the user sounds, in opposite directions at the two poles. The instrument is one word, one dollar, and judge-free; run per release, it reads the field’s anti-sycophancy training directly off model behavior.

[NLP-40] Understanding Machine Unlearning Through the Lens of Mode Connectivity

【速读】: 该论文旨在解决机器遗忘(Machine Unlearning)中对模型损失曲面结构与优化几何特性理解不足的问题,尤其关注如何在不进行从头训练的前提下有效移除模型中的特定信息。其核心解决方案是引入“遗忘模式连通性”(Mode Connectivity in Unlearning, MCU),通过分析独立训练的模型在参数空间中是否可通过低损失平滑路径相连,揭示未学习模型的分布特性。研究发现,许多已遗忘模型位于具有平滑保留/遗忘行为的连通基域内,而训练动态的变化会导致解进入不同的基域;同时,同一基域内的模型在隐私度量上可能差异显著,且遗忘过程呈现非线性特征。此外,线性连通性表明,大多数近似遗忘方法在机制上与重新训练存在本质区别。基于MCU的集成策略可提升模型泛化能力与抗重学习攻击鲁棒性,且MCU平滑性与遗忘难度呈正相关。该研究首次从模式连通性的视角系统剖析了机器遗忘的内在机理。

链接: https://arxiv.org/abs/2607.23970
作者: Jiali Cheng,Hadi Amiri
机构: University of Massachusetts Lowell (马萨诸塞大学洛厄尔分校)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: COLM 2026

点击查看摘要

Abstract:Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss landscape and optimization geometry of unlearning are poorly understood. In this paper, we study machine unlearning through the lens of mode connectivity–the phenomenon that independently trained models can often be connected by smooth low-loss paths in parameter space. We introduce \em mode connectivity in unlearning (MCU) and evaluate it across a range of settings, including curriculum learning, second-order optimization, and connectivity across different unlearning methods. We find that many unlearned models lie in connected basins with smooth retain/forget behavior, while changes in training dynamics can move solutions into different basins. MCU also reveals that models within the same basin can differ substantially on privacy metrics, and that unlearning progresses nonlinearly from the original model to the unlearned model. In addition, linear connectivity suggests that most approximate unlearning methods are mechanistically distinct from retraining. Finally, MCU-based ensembling can improve generalization and robustness to relearning attacks, and MCU smoothness correlates with unlearning difficulty. To our knowledge, this is the first study of machine unlearning through the lens of mode connectivity.

[NLP-41] Reality Monitoring in Large Language Models : Self-Knowledge That Transforms with Conversation Memory

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在对话中缺乏自我来源辨识能力的问题,即无法区分自身生成内容与用户输入信息,从而导致将自身错误当作外部事实接受。这一能力在人类认知中被称为现实监控(reality monitoring),其功能失效与幻觉、妄想和虚构等现象密切相关。研究通过两项实验及六种不同规模的LLM验证发现,模型在源归属判断上的表现高度依赖于对话记忆的结构设计:在记忆负担较轻时,模型对自生成内容的识别准确率接近上限;但一旦引入情景延迟(episodic delay)打破该捷径,准确率骤降,转而表现出对外部信息的脆弱优势。此外,反馈机制揭示了两类关键缺陷——部分模型出现内部与外部判断的混淆,另一些模型虽准确率提升但自信程度与正确性脱钩,此类解耦现象无法被现有基准检测。跨模型分析表明,问题根源在于模型的主动处理机制而非参数总量的累积效应。因此,随着AI系统承担更多自主、多轮交互任务,仅评估其知识掌握程度已不足以保障可靠性,追踪知识来源的可追溯性同样至关重要。

链接: https://arxiv.org/abs/2607.23927
作者: Saurabh Ranjan,Konstantina Sokratous,Brian Odegaard
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Neurons and Cognition (q-bio.NC)
备注:

点击查看摘要

Abstract:A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut. Feedback exposes two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. Across models, this pattern implicates active, not aggregate, parameter count. This suggests that as AI systems take on autonomous, multi-turn roles, evaluating what they know is not enough: tracking where that knowledge came from may matter equally.

[NLP-42] Understanding Tone-Dependent Inference Cost in Large Language Models

【速读】: 该论文旨在解决提示语语气(prompt tone)对大语言模型(Large Language Model, LLM)回答准确性与推理成本之间权衡的影响问题。其核心解决方案的关键在于揭示不同提示语气(从阿谀奉承到威胁性语气共七种)对模型输出令牌(output token)消耗量的显著影响,发现输出令牌长度的变化幅度远超准确率的变化,且在不同模型中,语气差异导致的推理成本波动最高可达44.3%。研究进一步分析了答案准确性与推理过程中平均输出令牌长度之间的帕累托最优前沿关系,发现对于ChatGPT 4o和5-nano模型,粗鲁语气表现更优;而对于Gemini 2.5 Flash及Lite模型,粗鲁与中立语气位于帕累托最优边界上。这表明提示语气不仅影响生成答案的质量,还显著影响可计费的推理资源消耗,为高效、低成本的LLM应用提供了关键优化方向。

链接: https://arxiv.org/abs/2607.23915
作者: Akhil Kumar,Om Dobariya
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 16 pages, 1 figure, 9-page Appendix

点击查看摘要

Abstract:We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 Question MMLU dataset for LLM models prompted in seven different tones from sycophantic to threatening. Our results show that the output-token-length variation substantially exceeded accuracy variation across all models. Output-token consumption varied by up to 44.3% across tone conditions. We also analyzed the tradeoff between the accuracy of the answers and the average output token length in the reasoning process. For the ChatGPT models 4o and 5-nano, the rude tone is quite dominant. For the Gemini models 2.5 Flash and 2.5 Flash Lite, the rude and neutral tones are dominant on the Pareto-optimal frontier. We find that prompt tone influences not only answer quality but also the amount of billable inference resources consumed by modern LLMs.

[NLP-43] riShieldRAG : A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation

【速读】: 该论文旨在解决生成式 AI(Generative AI)在检索增强生成(Retrieval-Augmented Generation, RAG)系统中面临的对抗性数据投毒攻击问题。具体而言,当外部知识库允许多方写入时,攻击者仅需注入少量精心构造的恶意文档,即可在查询时诱导大语言模型输出错误答案,严重威胁系统的可信性与安全性。现有单一防御机制(如困惑度过滤、查询重述、知识库扩展)均难以有效应对此类攻击,攻击成功率仍高达30%以上。为此,本文提出TriShieldRAG框架,其核心解决方案在于构建三个独立且形式化定义的防护层:数据摄入防护层(Ingest Guard),通过检测词汇与统计异常特征识别潜在投毒文档;检索重排序层(Retrieval Scorer),基于来源可信度与内容一致性加权计算信任分数,对检索结果进行动态重排序;以及跨大模型共识层(Cross-LLM Consensus),利用三类架构差异显著的语言模型(Claude、Mistral Small、Llama 3.2)进行联合推理,并允许一次有限的再检索以处理分歧。该方案依赖于“少数投毒”与“显式溯源标签”两个假设,实验证明,在面对原始PoisonedRAG中的非自适应攻击者时,该系统可将攻击成功率从约91%降至约13%,同时保持良性查询的准确率,显著提升了RAG系统在开放或半开放知识环境下的鲁棒性与可信度。

链接: https://arxiv.org/abs/2607.23838
作者: Susil Kumar Mohanty,Rohit Patel,Kosuru Yuvaraj,Jeenal Chaudhary,Disha Singhania
机构: 未知
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query time. This makes RAG useful for private data, fast-changing information, and reducing hallucination, but it also means the model’s answer is only as trustworthy as whatever the retriever hands it. If the knowledge base accepts writes from more than one party, an attacker needs only a handful of adversarial documents to steer the model toward a chosen wrong answer. PoisonedRAG demonstrated this: as few as five crafted documents flip an undefended system’s answer roughly 90% of the time, and three natural single-stage defenses (perplexity filtering, query paraphrasing, knowledge-base expansion) leave attack success at 30% or higher. We built TriShieldRAG to close that gap. Rather than relying on one checkpoint, we place three independent, formally specified rings across the pipeline: an Ingest Guard that screens documents for lexical and statistical poisoning signatures; a Retrieval Scorer that re-ranks the retrieved set by a provenance and consistency-weighted trust score; and a Cross-LLM Consensus stage that polls three architecturally diverse language models (Claude, Mistral Small, Llama 3.2) and allows one bounded re-retrieval on disagreement. We derive the conditions under which Rings 2 and 3 are expected to work: a minority-poison assumption and an explicit provenance-tag assumption. Our reported configuration is consistent with this analysis, though we have not yet run the controlled poison-fraction sweep needed to confirm it independently. Evaluated against the non-adaptive attacker from the original PoisonedRAG, over a 5,000-document Wikipedia knowledge base with 10 target questions, the full pipeline reduces attack success rate from roughly 91% to roughly 13% while preserving accuracy on benign queries. Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) Cite as: arXiv:2607.23838 [cs.CR] (or arXiv:2607.23838v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2607.23838 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-44] Latent-LoRA: Compact Latent-Space Adapters with Gradient-Free Routing for Continual Learning

【速读】: 该论文旨在解决大语言模型在持续学习(continual learning)过程中面临的灾难性遗忘问题,尤其针对现有基于低秩适应(LoRA)的持续学习方法中存在任务标识依赖或全局适配器叠加导致输出干扰的缺陷。其核心解决方案在于提出一种无需可训练路由模块、无须重放(replay-free)且参数高效的新框架——Latent-LoRA。关键创新点包括:首先,利用冻结预训练模型嵌入层的池化标记嵌入(pooled token embeddings),通过非梯度优化的高斯混合模型(Gaussian Mixture Model, GMM)实现无需任务身份信息的任务无关适配器选择;其次,通过奇异值分解(SVD)将每个任务的参数约束于预训练权重的主子空间内,实现紧凑的隐空间参数化,并结合正交正则化直接抑制任务间干扰。该方法不仅避免了可训练门控模块带来的遗忘风险,还显著降低了每任务参数量,实验在五个模型规模及两个主流持续学习基准上均达到当前最优性能,几乎实现零遗忘。

链接: https://arxiv.org/abs/2607.23837
作者: Reza Rahimi Azghan,Gautham Krishna Gudur,Giulia Pedrielli,Pavan Turaga,Hassan Ghasemzadeh
机构: Arizona State University (亚利桑那州立大学); The University of Texas at Austin (德克萨斯大学奥斯汀分校)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models generalize well to individual tasks but lack an inherent mechanism for learning them sequentially, leading to catastrophic forgetting. To mitigate this, LoRA-based continual learning methods allocate a separate low-rank adapter per task, yet existing approaches either require task identity at inference or sum all adapters indiscriminately, letting irrelevant branches distort the output. Recent gating-based solutions route inputs to the correct adapter but introduce trainable parameters that themselves need protection against forgetting. In this work, we observe that pooled token embeddings from a frozen LLM embedding layer already separate task distributions throughout the learning sequence. A Gaussian mixture model fitted on these embeddings, without any gradient-based training, is sufficient for task-agnostic adapter selection at test time. This eliminates the need for a learned gating module. On the adapter side, constraining each task’s parameters to the principal subspace of the pretrained weights via SVD yields a compact latent-space parameterization. Within this subspace, orthogonal regularization directly controls inter-task interference. The resulting system, Latent-LoRA, is replay-free, requires no trainable routing component, and uses substantially fewer parameters per task. Experiments across five model scales and two established continual learning benchmarks show state-of-the-art performance with near-zero forgetting.

[NLP-45] Kalypso: Relational LLM Serving

【速读】: 该论文旨在解决现有语义查询处理系统中大型语言模型(Large Language Models, LLM)服务机制缺乏对查询计划感知的问题,导致无法充分利用性能优化机会。当前基于请求中心的LLM服务系统独立处理每个查询,忽视了查询内部各语义操作符之间的结构关系,从而造成计算资源浪费。其核心解决方案是提出“关系型LLM服务”(relational LLM serving)这一新抽象,使LLM服务能够感知语义查询的结构,并在保持查询语义一致性和输出准确性的前提下,实现跨语义操作符的流水线执行。关键创新在于:当中间元组从一个操作符直接传递至下一个操作符时,可复用其键值缓存(KV-cache)状态,避免重复计算。为此,论文设计并实现了Kalypso系统,通过提供语义查询计划的API接口,采用自适应、内存感知的调度算法执行查询。该系统面临一种新型在线调度问题,需在流水线执行与GPU内存压力管理之间进行权衡,以在缓存被驱逐前有效重用KV-cache状态。其调度器动态调整内存分配,在上游并行度、下游进度和GPU利用率之间实现平衡。实验结果表明,相较于传统的请求中心式LLM服务,Kalypso在多种工作负载下显著缩短了查询完成时间,最高加速比达4.57倍,验证了查询感知的LLM服务在提升语义查询执行效率方面的巨大潜力。

链接: https://arxiv.org/abs/2607.23815
作者: Hojae Son,Md Ashraful Islam,Huy Gia Cao,Hui Guan,Marco Serafini
机构: University of Massachusetts Amherst (马萨诸塞大学阿默斯特分校)
类目: Databases (cs.DB); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 14 pages, 12 figures

点击查看摘要

Abstract:Large language models are increasingly used as semantic operators for filtering, extracting, ranking, joining, and transforming unstructured data. Existing semantic query processing systems invoke request-centric LLM serving systems that are unaware of the query plan, leaving substantial performance opportunities unused. This paper introduces relational LLM serving, an abstraction that makes LLM serving aware of semantic query structure while preserving query semantics and output accuracy. The key opportunity is pipelined execution across semantic operators: when intermediate tuples flow directly from one operator to the next, their KV-cache state can be reused instead of recomputed. We present Kalypso, a relational LLM serving system that exposes an API for semantic query plans and executes them using an adaptive, memory-aware scheduling algorithm. Kalypso addresses a new online scheduling problem in which pipelined operator execution is coupled with GPU memory pressure management to reuse KV-cache state in the serving engine before eviction. Its scheduler continuously adjusts memory allocations to balance upstream parallelism, downstream progress, and GPU utilization. Our evaluation shows that Kalypso improves query completion time over baselines using request-centric LLM serving, with speedups up to 4.57x across diverse workloads, demonstrating that query-aware LLM serving can substantially improve the efficiency of semantic query execution. Comments: 14 pages, 12 figures Subjects: Databases (cs.DB); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) Cite as: arXiv:2607.23815 [cs.DB] (or arXiv:2607.23815v1 [cs.DB] for this version) https://doi.org/10.48550/arXiv.2607.23815 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-46] Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

【速读】: 该论文旨在解决金融领域中真实场景下英文财报电话会议(earnings calls)的自动语音识别(ASR)评估难题。现有基准在数据覆盖度、行业平衡性及细粒度评估能力方面存在不足,难以全面反映模型在复杂真实场景中的表现。为此,论文提出Earnings25,一个面向金融领域的ASR评测基准,其关键创新在于构建了两个互补的测试集:testset-full(498小时完整的标普500公司2025年第四季度财报电话会议录音)与testset-segmented(46小时按行业均衡采样的290个片段),并提供对齐的转录文本与结构化元数据(包括发言者角色、行业标签及通话结构),支持基于发言人和行业的精细化评估,突破传统仅依赖平均词错误率(WER)的局限。解决方案的核心在于通过真实、大规模且具备标注信息的数据集,实现更可复现、更具场景适应性的模型评估。

链接: https://arxiv.org/abs/2607.23813
作者: Denglin Jiang,Haoran Zhou,Anshul Wadhawan,Brendan Fahy,Vinay Ramesh,David Weisberg,Dmitriy Derkachevskiy,Helen Sheehan,Srivas Prasad,Michele Franceschini
机构: Bloomberg(彭博社)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 5 pages, 0 figures, 5 tables

点击查看摘要

Abstract:We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full English-language SP 500 earnings calls from Q4 2025, and (ii) testset-segmented, a 46-hour industry-balanced set of 290 segments sampled from English-language U.S. earnings calls in 2025. The benchmark provides aligned transcripts and structured metadata, including speaker roles, industry labels, and call structure, enabling speaker- and industry-aware evaluation beyond aggregate word error rate (WER). We report reproducible baselines for Whisper and Parakeet-TDT using standardized scoring.

[NLP-47] Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers ICASSP

【速读】: 该论文旨在解决在资源受限的设备端实现高保真、实时语音合成的挑战,特别是针对苹果矩阵协处理器(AMX)的严格计算与内存预算。其核心问题是:如何在不牺牲语音质量的前提下,高效地将由基础模型生成的语义音频标记(semantic audio tokens)转换为高质量音频,同时保持极低的内存占用和恒定的计算开销。解决方案的关键在于提出一种记忆高效的解标记器(detokenizer)架构,采用三组件设计——流式编码器、时序解码器与深度解码器,系统性解耦时间维度与深度维度的处理。该架构通过一个可复用的深度解码器,结合基于扩散变压器(DiT)风格的阶段条件控制,实现所有残差向量量化(RVQ)层级的自回归生成,取代传统多解码器架构中每个层级独立解码的设计;同时引入因果滑动窗口注意力与固定窗口键值缓存机制,使内存复杂度恒定且不随序列长度增长。该方案在AMX上实现了每生成步骤约10毫秒的延迟(约为实时速度的16倍),峰值运行内存仅21 MB,本地存储占用329 MB,支持长达20至320秒的连续语音流合成。相比传统基于Transformer或GAN的方法所呈现的线性与二次级内存增长,该架构实现了恒定的小型化内存足迹。消融实验验证了各组件的有效性,音频质量评估表明,在保持高保真度的同时,相较先前的设备端文本到语音系统,整体平均意见得分(MOS)提升+0.28(4.15 vs. 3.87),对话语音子集提升+0.42(4.24 vs. 3.82),在10亿参数激活规模下显著提升了语音自然度与表现力。

链接: https://arxiv.org/abs/2607.23811
作者: Dongseong Hwang,Prasanth Yadla,Kaan Elgin,Shifas Padinjaru Veettil,Sivanand Achanta,Dipjyoti Paul,Ramya Rasipuram,Tyler Johnson,Emad Soroush,Chung-Cheng Chiu,Zhifeng Chen
机构: Apple(苹果)
类目: ound (cs.SD); Computation and Language (cs.CL)
备注: 11 pages, ICASSP

点击查看摘要

Abstract:Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple’s most powerful on-device foundation model. This work presents the memory-efficient audio synthesis architecture behind that capability: a detokenizer that converts the semantic audio tokens emitted by the foundation model into high-fidelity audio within the tight compute and memory budget of the Apple Matrix Coprocessor (AMX). We convert semantic audio tokens to a residual vector quantization (RVQ) representation with a three-component design, a streaming encoder, a temporal decoder, and a depth decoder, that systematically decouples temporal and depth processing. A single reusable depth decoder with Diffusion Transformer (DiT)-style stage conditioning generates all RVQ levels autoregressively, replacing the dedicated per-level decoders of prior multi-decoder architectures, while causal sliding window attention with fixed-window key-value caching yields constant memory complexity independent of sequence length. Deployed on the AMX, the detokenizer sustains roughly 10 ms per generation step, about 16x faster than real time, with a peak runtime memory of only 21 MB and 329 MB of on-device assets, enabling continuous streaming synthesis of 20-320 seconds of audio. This constant, small footprint replaces the linear and quadratic memory scaling of conventional transformer- and GAN-based approaches. Ablation studies validate the key architectural components, and audio quality assessment confirms that the architecture maintains synthesis fidelity while achieving efficiency gains over existing methods. Operating at a 1-billion-parameter activation size within AFM 3 Core Advanced, it improves Mean Opinion Score by +0.28 overall (4.15 vs. 3.87) and by +0.42 on conversational speech (4.24 vs. 3.82) over the prior on-device text-to-speech system.

[NLP-48] Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages INTERSPEECH2026

【速读】: 该论文旨在解决印度多语言环境下语音技术中缺乏高质量、覆盖广泛语言的说话人分离(speaker diarization)与自动语音识别(ASR)联合基准数据集的问题。当前主流语音系统在处理印度语境下的复杂口语特征(如英语混用、方言差异及频繁的说话人重叠)时表现有限,且现有资源大多集中于少数语言或特定场景。其解决方案的关键在于构建Indic DiarBench——一个涵盖印度22种官方语言的开放性基准数据集,包含近108小时来自近场会议、远场录音及真实环境中的多说话人音频,并配有经人工校正的时间对齐说话人标注转写文本。该数据集特别强调捕捉印度语境下的对话细微特征,为评估和推动多语言、包容性语音技术的发展提供了坚实基础。

链接: https://arxiv.org/abs/2607.23808
作者: Deovrat Mehendale,Aditya Mehndiratta,Dhruv Rathi,Kaushal Bhogale,Mitesh M. Khapra
机构: Sarvam AI(萨尔瓦姆人工智能); AI4Bharat, IIT Madras(印度理工学院马德拉斯分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 5 pages, 2 figures, Interspeech 2026 conference

点击查看摘要

Abstract:In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to advance inclusive, multilingual speech technology research for Indian languages.

[NLP-49] How Context Attribution Handles What the Model Already Knows

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在上下文归因(context attribution)中无法有效区分上下文学习(In-Context Learning, ICL)与权重内知识(In-Weight, IW)贡献的问题。当输入上下文与模型训练数据存在重叠时,现有归因方法会混淆ICL与IW的贡献,导致归因得分不可靠。其解决方案的关键在于:1)提出一套基于四个新指标(基础模型上下文归因得分,BCS;跨模型上下文归因一致性,CAC;归因保真度得分,APS;源分离精度,SSP)的评估协议;2)构建一个带有真实来源标签的基准数据集(WMDP-Cyber++),以系统性地评估在IW重叠场景下的归因性能。实验结果表明,四种主流归因方法在存在IW重叠时均产生不忠实的归因结果;进一步地,尽管尝试通过贡献分数对ICL与IW进行源分离,但这些方法仍无法实现有效解耦,揭示了当前归因机制的根本局限性。

链接: https://arxiv.org/abs/2607.23804
作者: Quoc-Huy Trinh,Lin Zhu,Sebastian Szyller
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works show the initial success in attributing the con- tributive score of the contexts. However, we observe that when the context overlaps with the training data, these methods can- not disentangle in-context from in-weight (IW) contributions, producing unreliable scores. Based on this observation, in this work, we introduce: 1) an evaluation protocol that relies on four new metrics (base-model context attribution score (BCS), cross-model context attribution consistency (CAC), attribution preservation score (APS), source separation pre- cision (SSP)) and 2) a benchmark dataset (WMDP-Cyber++) with ground-truth provenance labels to systematically assess attribution under IW overlap. In our experiments across four well-known context attribution methods, we demonstrate that they provide unfaithful attribution when the knowledge from the context also exists in the weights. Finally, we adapt these methods for source separation (IW vs. in-context learning (ICL)) and show that they cannot do the disentanglement based on the contributive score

[NLP-50] Zing: Social Mind for LLM s

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在长期人机交互环境中缺乏社会智能(Social Intelligence)的问题,即模型难以准确推断心理状态、追踪社会关系、推理规范性行为并根据上下文动态调整自身行为。其核心挑战在于现有模型在社会认知能力上存在显著短板,且缺乏系统性的评估、内化与部署机制。解决方案的关键在于提出一个集成框架Zhijing,包含三个协同模块:首先,构建了基于心理学理论的SoMBench基准,覆盖3个主维度、17个次级维度及71种任务范式,通过标准化的场景设计与专家验证实例,实现对社会智能的精准测量;其次,提出Zing训练方案,融合监督微调、在线策略蒸馏与基于评分标准的强化学习,实现社会智能的参数化内化,显著提升模型在五项社会认知基准上的表现;最后,设计Actio推理架构,引入四类受控运行时支持——PRISM(程序化引导)、Starling(实时心理状态表征)、SAGE(可复用经验库)和门控检索增强生成(gated RAG,用于外部社会与规范知识),实现部署阶段的社会智能动态增强。实证表明,该框架在多个基线模型与评测集上均取得显著性能提升,验证了评估、内化与运行时支撑三者协同对构建真正具备社会智能的LLMs的重要性。

链接: https://arxiv.org/abs/2607.23740
作者: Zing Team,Ao Xiang,Bi Jingping,Chen Jiahui,Chen Lehan,Chen Yilin,Cheng Xueqi,Fan Yixing,Gan Kairong,Gao Haowen,Gao Jinhua,Gao Shuxuan,Gong Chang,Guo Jiafeng,Guo Ruijie,Han Zhouyu,He Guangfu,He Yichun,Jiang Shuo,Jing Shaoling,Jing Ya,Lei Chenhao,Lei Yan,Li Anqi,Li Chengao,Li Haoyu,Li Shitian,Liang Xinjian,Liu Zhaoge,Lyu Xingyu,Nie Zhuwei,Pang Liang,Quan Zeping,Shan Shiguang,Shen Huawei,Tang Xinran,Tian Feng,Wang Qian,Wang Ruiping,Wang Xiaohong,Xia Zaiyu,Xiao Yi,Xu Jiayuan,Xu Kehan,Xu Qianqian,Xu Tianyu,Xu Yongjun,Yang Haoming,Yang Jun,Yao Di,Yu Xiaoming,Zhang Futong,Zhang Jie,Zhang Shixuan,Zhang Yuxuan,Zhao Xinyu,Zhao Zhuoran,Zhong Yunfei,Zhu Shengyu
机构: Zing Team
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context. This report presents Zhijing, an integrated framework for measuring, internalizing, and grounding social intelligence. For measurement, we introduce SoMBench, a psychology-grounded benchmark spanning 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms. It controls question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Evaluation of 20 representative LLMs reveals substantial headroom: the best model achieves only 72.08% overall accuracy, and none of the 17 secondary dimensions reaches the 90% near-ceiling band. For internalization, we develop Zing, a diagnosis-driven training recipe combining supervised fine-tuning, on-policy distillation, and rubric-based reinforcement learning. Across five social-cognition benchmarks, Zing consistently outperforms its base models, with Zing-27B-Stage2 achieving the best average score and Zing-32B-Stage2 remaining competitive with DeepSeek-V4-Pro. For deployment-time grounding, we build Actio, a harness-controlled inference architecture that routes four typed supports into reasoning: PRISM for procedural guidance, Starling for runtime mental-state representation, SAGE for reusable experience, and gated RAG for external social and normative knowledge. Across five base models and three benchmarks, the full harness improves 14 of 15 model-benchmark pairs and is best or tied for best in 8, demonstrating the effectiveness of typed runtime support. Together, these results show that socially intelligent LLMs require coordinated advances in evaluation, parametric internalization, and deployment-time grounding.

[NLP-51] Formally Verified Synthesizable Floating-Point Data Types in ARCH HDL

【速读】: 该论文旨在解决在由语言模型生成的硬件描述语言ARCH中,实现符合IEEE-754标准的二进制32位(FP32)与脑浮点16位(bfloat16,BF16)算术运算的可验证性与正确性问题。其核心挑战在于确保算术单元在不同形式化表示之间的一致性,尤其是在面对高复杂度乘法操作时的可证明正确性。解决方案的关键在于构建一个统一的比特向量中间表示(bit-vector IR),所有算术操作(包括比较、转换、加减、乘法及融合乘加,FMA)仅需定义一次,并通过同一源代码生成三种可执行且可验证的产物:可综合的SystemVerilog代码、SMT-LIB逻辑模型以及Lean 4形式化证明模型。三者结构保持一致,通过Yosys到SMT的比对器(miter)自动验证SystemVerilog与SMT模型等价性。对于乘法无关的操作(如比较、加减、转换及所有二进制BF16运算),采用穷尽式验证,证明其与SMT-LIB浮点理论完全等价;而对于乘法相关的高难度操作(如FP32乘法和FMA),则在Lean 4中以值级“四舍六入五成双”(round-to-nearest-even)规范进行无依赖(sorry-free)的形式化证明,覆盖所有2^96个输入。为应对物理特性中的关键路径瓶颈——精确宽470位的数据通路无法流水化,研究者重构了FMA模块为98位的带保护/舍入/粘滞(guard/round/sticky)流水结构,在Nangate45工艺下实现268 MHz运行频率,并通过形式化证明确认其与原始精确版本在所有输入下比特完全一致,从而继承其正确的舍入属性。该等价性得以实现的核心原因在于双方共享相同乘法器,可在证明中直接消去,避免了求解乘法等价性的计算难题。所有机器可检查的断言均锚定于一个带标签的开源发布版本,保障可复现性与可信度。

链接: https://arxiv.org/abs/2607.23715
作者: Shuqing Zhao
机构: 未知
类目: Computation and Language (cs.CL); Programming Languages (cs.PL)
备注: 8 pages, 2 figures

点击查看摘要

Abstract:We report the design and end-to-end verification of first-class IEEE-754 binary32 (FP32) and bfloat16 (BF16) arithmetic for ARCH, a hardware description language intended to be generated by language models. Every operator - comparisons, conversions, add, sub, mul, and fused multiply-add (FMA) - is described once against a single bit-vector IR and rendered three ways from one source: synthesizable SystemVerilog, an SMT-LIB model, and a Lean 4 proof model. The three artifacts cannot drift apart structurally, and the residual per-node printer correspondence is machine-checked: a Yosys-to-SMT miter proves the emitted SystemVerilog equivalent to the SMT model for all 24 operators. Verification splits at the solver-tractability frontier: multiplier-free operators (comparisons, add/sub over all 2^64 inputs, conversions, and all binary BF16 arithmetic) are proved exhaustively equivalent to the SMT-LIB FloatingPoint theory; the SAT-hard multiplier-bearing operators (FP32 mul and FMA) are proved correctly rounded in Lean, sorry-free, against a value-level round-to-nearest-even specification over exact dyadic values. Physical characterization exposed the FMA as the timing outlier: its exact-wide 470-bit datapath does not pipeline in our flow. We reimplemented it as a bounded 98-bit guard/round/sticky datapath that pipelines to 268 MHz on Nangate45, and proved, in Lean and over all 2^96 inputs, that it is bit-identical to the exact-wide reference, so it inherits the reference’s proven correct rounding. The equivalence is tractable precisely because the shared multiplier appears on both sides and cancels: neither a SAT solver nor the proof ever solves a multiplier equivalence. (The BF16 FMA is deliberately an FP32-accumulating fusion, characterized as exactly that.) All machine-checked claims are pinned to a tagged open-source release.

[NLP-52] An empirical investigation into the properties of standard word embeddings

【速读】: 该论文旨在解决自然语言处理中词嵌入(Word Embedding)表示方法的多样化与实践选择难题,即如何有效计算和应用连续向量空间中的词表示以提升下游任务性能。其核心问题在于现有多种词嵌入模型(如Word2Vec、GloVe、FastText等)在训练机制、语义表达能力及适用场景上存在差异,导致研究者与开发者在实际应用中面临工具选型与性能评估的挑战。解决方案的关键在于系统性地回顾主流词嵌入生成机制,对比分析公开可用的工具包(如gensim、spaCy)与预训练嵌入矩阵的特性,并通过实验验证典型实现的效果,从而揭示不同方法在语义相似性、上下文敏感性及泛化能力等方面的优劣,为实际应用提供可参考的技术选型依据。

链接: https://arxiv.org/abs/2607.23675
作者: Salomon Kabongo
机构: African Institute for Mathematical Sciences (AIMS); North-West University
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: African Institute for Mathematical Sciences (AIMS) - South Africa, University of the Western Cape

点击查看摘要

Abstract:The embedding of word sequences into continuous vector spaces has been one of the most important developments in Natural Language Processing in the recent past. Such embeddings have found application in areas such as Automatic Speech Recognition, Machine Translation, Sentiment Analysis and many more. This essay reviews the various mechanisms that have been proposed for the calculation of word embeddings, investigates popular toolkits and embedding matrices that are available in the public domain, and experiments with one or more selected implementations to better understand their characteristics. La représentation vectorielle continue de mots a été l’un des développements les plus importants dans le domaine du traitement automatique du langage naturel au cours des dernières années. Ces représentations ont trouvé application dans des domaines tels que la reconnaissance vocale, la traduction automatique, l’analyse des sentiments, etc. Ce travail passe en revue les différents mécanismes proposés pour le calcul de ces vecteurs de mots, étudie les kits d’outils populaires et les matrices disponibles publiquement en ligne, et expérimente avec une ou plusieurs implémentations sélectionnées pour mieux comprendre leurs caractéristiques. Comments: African Institute for Mathematical Sciences (AIMS) - South Africa, University of the Western Cape Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.23675 [cs.CL] (or arXiv:2607.23675v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2607.23675 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-53] EmoTrace: An Emotion Trajectory-Centered Framework for Psychological Support Dialogue Generation

【速读】: 该论文旨在解决当前基于大语言模型(LLM)的心理咨询对话生成中,数据集存在求助者情绪状态过于稳定、情感动态变化有限以及对咨询师引导过度顺从等问题,导致模型难以有效应对情绪不稳定的实际咨询场景。此外,现有方法多聚焦于问题解决导向的回应,忽视了心理咨询中情感聚焦互动的重要性。为此,论文提出EmoTrace——一种以建模求助者情感轨迹为核心的多轮对话语料生成框架,其关键在于构建求助者的认知画像,引入包含情感模板(emotional schemas)及激活机制的求助者模块、咨询师模块与情感轨迹控制模块,从而实现求助者情感表达的多层次丰富化与咨询师共情回应的精准化。该设计显著提升了生成对话的情感丰富度与共情质量。

链接: https://arxiv.org/abs/2607.23648
作者: Kaitong Weng,Lixin Liu,Zihao Liu,Bo Wang,Shiguang Ni
机构: Shenzhen International Graduate School, Tsinghua University (清华大学深圳国际研究生院)
类目: Computation and Language (cs.CL)
备注: 36 pages, 20 figures

点击查看摘要

Abstract:Using large language models (LLMs) to assist psychological counseling is an important task in the field of natural language processing. The construction of high-quality psychological support dialogue corpora serves as a critical foundation for training counseling-oriented conversational models. However, existing data generation approaches generally suffer from several limitations, including emotionally stable seekers, limited variation in emotional dynamics, and a high degree of compliance with counselors’ guidance. These issues result in LLM that lack the capability to effectively respond to emotionally unstable scenarios. In addition, counselor responses are typically driven by problem-solving objectives, thereby overlooking the role of emotion-focused interaction, which are essential in psychological counseling. To address these gaps, we propose EmoTrace, a multi-turn dialogue corpus generation framework centered on modeling seekers’ emotional trajectories. we construct seekers’ cognitive profile and introduce a seeker module with emotional schemas and an associated activation mechanism, a counselor module, and an emotional trajectory control module, thereby enhancing the layering of the seeker’s emotional expression and the counselor’s targeted empathic expression. Experimental results demonstrate that the proposed method outperforms existing approaches in terms of emotional richness and empathy quality.

[NLP-54] Where Is the Cost of Third-Party API Routers in Agent ic Software Development?

【速读】: 该论文旨在解决第三方API路由器(third-party API router)在生成式AI驱动的代码代理(coding agent)工作流中可能引入的安全控制缺口问题。随着高自主性代码代理的广泛应用,其与上游大语言模型(LLM)提供商之间的通信路径被第三方路由器所占据,而该路由器具备对请求与响应的完全访问权限,却缺乏机制验证模型输出与最终执行的代码库级操作之间的一致性。这导致客户端权限控制机制可能失效,进而引发难以察觉的恶意行为。为探究此风险的实际影响,研究设计并实施了一项实证研究,系统评估了四种渐进式隐蔽攻击级别:响应替换(L1)、响应追加(L2)、LLM润色注入(L3)以及结合分布对齐的LLM润色注入(L4)。研究提出SIDEL框架,支持追踪、重放、注入与防御评估,并构建了包含400个样本的专用数据集。实验表明,路由器端干预可显著改变代码库层面的操作行为,且现有客户端防护手段几乎无法检测此类攻击;在未部署额外缓解措施的情况下,所有测试代码代理在各攻击层级上的防御成功率均为0%。尽管基于白名单的执行控制和大语言模型审查可提升一定抗性,但无法完全恢复端到端控制能力,从而凸显出由服务提供方保障输出完整性的必要性。

链接: https://arxiv.org/abs/2607.23624
作者: Donghao Fu,Jingxin Li,Xue Jiang,Yihong Dong
机构: Shanghai Jiao Tong University (上海交通大学); Beihang University (北京航空航天大学); Peking University (北京大学)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Third-party API routers have become a common layer that unifies access across increasingly diverse LLM providers. In coding-agent workflows, high-autonomy operation is widely adopted because it reduces interaction overhead. As a result, a third-party API router, which sits between the agent and the upstream provider, inevitably occupies the trusted path. It can inspect and modify every request and response, yet no mechanism verifies alignment between the provider’s output and the repository-level actions ultimately executed by the agent. Consequently, client-side permission mechanisms may become ineffective in practice. Whether this control gap produces real, hard-to-detect effects on software development tasks remains empirically unmeasured. In this paper, we conduct an empirical study of router-side injection in coding agents, examining four intervention levels of increasing subtlety: Response Substitution (L1), Response Append (L2), LLM-Polished Injection (L3), and LLM-Polished with Distribution Alignment Injection (L4). Moreover, we develop SIDEL, a framework for trace recording, replay, injection, and defense evaluation, with a curated dataset of 400 samples. We evaluate four representative coding agents, and further evaluate whitelist-based execution control and LLM review. Router-side intervention substantially alters repository-level actions and remains difficult for existing client-side safeguards to detect. Without additional mitigations, all evaluated agents achieved a defense success rate of 0 percent across all injection levels. Client-side mitigations and reactive reviews improve resistance but do not fully restore end-to-end control, motivating provider-side output-integrity guarantees. Our code is available at this https URL.

[NLP-55] GEMCo: A Validated Ethically Releasable Proxy for Inaccessible Counselling Data

【速读】: 该论文旨在解决隐私与伦理限制下难以获取真实心理咨询数据的问题,从而阻碍了相关语言研究的开展。其核心挑战在于如何在不泄露敏感信息的前提下,构建一个可公开使用的、具有代表性的替代数据集。解决方案的关键是提出GEMCo——一个由人工撰写的、可释放的代理数据集,包含86组完整的德语电子邮件心理咨询对话(共728条消息)、专家编撰的案例以及由受训演员扮演的咨询师会话。该数据集通过与保留的124个真实咨询对话进行对比,从咨询策略和来访者情绪两个维度验证其有效性,结果显示两者差距虽存在但较小,具备较高的保真度。研究进一步采用生成式验证方法增强分析可信度,且所提出的验证范式可推广至其他无法共享真实数据但可构建人工代理的领域。由于GEMCo在设计上不包含任何敏感信息,符合隐私保护要求,因此可被安全发布,为心理语言学与生成式AI在心理健康领域的研究提供了首个可用的开放资源。

链接: https://arxiv.org/abs/2607.23621
作者: Philipp Steigerwald,Eric Rudolph,Mara Stieler,Jennifer Burghardt,Jens Albrecht
机构: Technische Hochschule Nürnberg Georg Simon Ohm, Germany
类目: Computation and Language (cs.CL); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:This paper presents GEMCo, a releasable, human-written proxy for inaccessible counselling data: 86 complete German e-mail counselling conversations (728 messages), expert-authored cases and counsellor sessions with trained role-players. It is validated against a held-out reference of 124 real counselling conversations. The proxy and the real conversations are measured against each other in counsellor strategies and client emotions. The gap is detectable but small. A generative validation supports the analysis. The validation method itself generalises to any domain where real data cannot be shared but a human-made proxy can. Privacy and ethics keep real counselling data closed. GEMCo carries none by design and can be released – a first step toward language research in this domain.

[NLP-56] MS-GPT : Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model

【速读】: 该论文旨在解决从串联质谱(MS/MS)数据中实现无参考库依赖的分子结构从头解析(de novo structure elucidation)这一核心逆问题。现有方法多依赖于参考数据库或预定义候选集,而纯从头方法虽可直接由谱图生成结构,但其主流范式——先从谱图预测分子指纹(fingerprint),再基于指纹解码结构——导致训练与推理之间的不匹配:模型在训练时使用的是由真实分子精确计算出的“真值指纹”(oracle fingerprint),而在推理阶段却需处理由谱图噪声诱导的模糊后验指纹,通常被简化为单一阈值化指纹,从而损失信息并影响结构生成质量。本文提出MS-GPT,将指纹驱动的从头解析重构为对条件分子语言模型的谱图诱导后验查询问题。具体而言,MS-GPT以指纹和分子式为条件,构建一个条件分子语言模型,并通过主动位密度校准(active-bit density calibration)将谱图诱导的后验分布转化为围绕真值指纹流形的一组邻近指纹查询,显著提升后验表示的多样性与准确性。随后,在该指纹带内采样候选结构,并依据生成频率共识进行排序,实现高精度结构推荐。此外,引入轻量级LoRA适配器以缓解领域特异性后验偏差,同时保留预训练的分子先验知识。在NPLIB1和MassSpecGym数据集上,MS-GPT分别达到29.8%/41.1%和23.9%/28.7%的Top-1/Top-10精确匹配准确率,刷新当前最优性能。实验表明,通过扩大候选池规模,仅增加少量推理开销即可持续提升召回率,验证了其高效自回归分子生成的有效性。

链接: https://arxiv.org/abs/2607.23607
作者: Xin Zhao,Yumin Liu,Zhuo Li,Weichu Zheng,Feng Zhu,Xiaokang Yang,Yaohui Jin,Yanyan Xu
机构: Shanghai Jiao Tong University (上海交通大学); ByteDance Inc. (字节跳动公司)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 18 pages, 14 figures, and 9 tables, including appendices. Source code and model checkpoints are available at this https URL

点击查看摘要

Abstract:Molecular structure elucidation from tandem mass spectra (MS/MS) is a central inverse problem in analytical chemistry. Most existing approaches to MS/MS identification remain tied to reference libraries or predefined candidate sets, whereas de novo methods aim to generate structures directly from spectra. A common de novo route predicts a molecular fingerprint from the spectrum and then decodes structures from it, enabling decoder pretraining on large molecule-only corpora. However, this paradigm creates a training-inference mismatch: the decoder is trained on oracle fingerprints computed from molecules, but at inference it is queried with a noisy spectrum-induced fingerprint posterior that is typically collapsed to a single thresholded fingerprint. We introduce MS-GPT, which recasts fingerprint-mediated de novo elucidation as spectrum-induced posterior querying of a conditional molecule-language model. MS-GPT conditions a molecule-language model on fingerprints and formulas, then converts the spectrum-induced posterior into a band of fingerprint queries near the oracle-fingerprint manifold through active-bit density calibration. Candidates sampled across this band are pooled and ranked by generation-frequency consensus. A lightweight LoRA adapter further mitigates domain-specific posterior bias while preserving the pretrained molecular prior. On NPLIB1 and MassSpecGym, MS-GPT sets a new state of the art, reaching Top-1/Top-10 exact-match accuracy of 29.8%/41.1% and 23.9%/28.7%, respectively. Candidate-pool scaling shows that efficient autoregressive molecular generation continues to improve recall with a little additional inference cost. The source code and model checkpoints are available at this https URL.

[NLP-57] HiTMS: A High-Throughput Multi-Stream Linguistic Steganography Framework

【速读】: 该论文旨在解决生成式语言隐写(Generative Linguistic Steganography)中现有单流方案在批量多流推理场景下的两大局限:一是缺乏协议层面的多流支持,二是简单的批处理无法隐藏数据槽占用情况与载荷完成状态。其解决方案的关键在于提出HiTMS框架,通过在多轮交互中将秘密信息分布于多个响应之中,每轮在单次批处理调用内嵌入并提取多个数据流,从而分摊模型调用开销,显著提升吞吐量。为确保信息可恢复性,HiTMS采用自描述帧结构,并基于密钥导出的调度机制绑定流与槽位,同时以伪数据填充空槽,实现精确恢复的同时有效隐藏活跃流的数量。该框架与具体语言模型及隐写编码器无关,在八组不同数据集-模型-编码器组合下,八流配置的HiTMS相较单流基线最高提升4.3倍的嵌入与提取速度,且平均使隐写分析器的AUROC从0.681降至0.601;进一步实验表明,随着并发流数从4增至64,吞吐量优势持续保持。

链接: https://arxiv.org/abs/2607.23597
作者: Ruiyi Yan,Yugo Murawaki,Zhongliang Yang
机构: Kyoto University(京都大学); Beijing University of Posts and Telecommunications(北京邮电大学)
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Generative linguistic steganography conceals secret bits within the sampling randomness of large language models. Existing schemes are single-stream, conveying an entire secret through a single response to a single prompt. This convention incurs two limitations: it provides no protocol-level support for batched multi-stream inference, and naive co-batching does not conceal slot occupancy or payload completion. We propose HiTMS, which distributes a secret across multiple responses produced jointly over successive rounds of interaction. Each round embeds and extracts several streams within a single batched call, thereby amortizing the cost of model invocation and substantially improving throughput. To ensure recoverability, HiTMS wraps each response in a self-describing frame and employs a key-derived schedule that binds streams to slots and fills unused slots with decoys, guaranteeing exact recovery while concealing the number of active streams. The framework is agnostic to both the language model and the steganographic coder. Across eight dataset-model-coder settings, eight-stream HiTMS achieves up to 4.3 times higher embedding and extraction speeds than single-stream baselines, while reducing the steganalyzer AUROC from 0.681 to 0.601 on average. Additional experiments with 4 to 64 streams demonstrate sustained throughput gains as concurrency increases. GitHub repository for this work is this https URL.

[NLP-58] Language Shapes Instruction Hierarchy Compliance in Multilingual LLM s

【速读】: 该论文旨在解决多语言环境下指令优先级(Instruction Hierarchy, IH)合规性不稳定的问题,尤其关注现有评估体系长期局限于英语语境,导致对多语言场景下指令优先级机制的可靠性缺乏充分认知。其核心解决方案是提出XIHBench基准,涵盖六种语言、四个领域及三种指令优先级设置,系统评估同语言与跨语言冲突下的IH表现。研究发现两个关键现象:一是指令优先级的合规性呈现显著的语言依赖性不对称——高优先级位置中某语言增强合规性时,低优先级位置反而可能破坏整体控制;二是跨语言冲突下的合规性普遍高于同语言冲突,称为“语言边界效应”(Language Boundary Effect);此外,模型对特定语言的偏好会强化其在低优先级时的不可覆盖性,从而引发多语言环境下的可靠性与安全风险。这些发现揭示了当前生成式AI在多语言部署中潜在的控制漏洞,强调需构建更具鲁棒性的跨语言指令治理机制。

链接: https://arxiv.org/abs/2607.23545
作者: Jiwon Moon,Yerin Hwang,Kyomin Jung
机构: Seoul National University (首尔国立大学)
类目: Computation and Language (cs.CL)
备注: Code and data are available at this https URL

点击查看摘要

Abstract:Instruction hierarchy (IH) requires models to prioritize instructions by source, ensuring that higher-priority instructions override lower-priority ones. Despite its importance for safe and controllable deployment, existing evaluations have focused almost exclusively on English, leaving it unclear whether IH compliance remains stable in multilingual settings. We introduce XIH-Bench, a benchmark for multilingual IH evaluation with both same-language and cross-language conflicts across six languages, four domains, and three IH settings. Across models, we find two consistent patterns. First, IH compliance exhibits a clear language-dependent asymmetry: a language that strengthens compliance in the higher-priority position can become disruptive in the lower-priority position. Second, cross-language conflicts yield higher compliance than same-language conflicts, a phenomenon we term the Language Boundary Effect. We further show that language specialization can make lower-priority instructions in model-favored languages harder to override, creating multilingual reliability and security risks.

[NLP-59] Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration

【速读】: 该论文旨在解决生成式 AI 在低资源语言中生成共情性心理健康咨询回复能力不足的问题。其核心挑战在于,现有大语言模型(LLM)在非主流语言环境下缺乏对文化语境、情感敏感度及伦理规范的准确理解与响应能力。解决方案的关键在于提出一种任务特定的提示工程框架——角色扮演反思链式思维咨询框架(RP-RCAF),该框架通过融合专家撰写的少量示范案例与结构化自我反思机制,引导模型以同理心顾问角色生成支持性强、文化适切且符合伦理的心理咨询回应。同时,研究构建了基于Grok 4的响应评估与评分框架(G-REFS),结合自动化评估与临床心理学专家验证,从情感敏感性、文化适宜性、语言清晰度和伦理合理性四个维度实现多维评价。实验结果表明,RP-RCAF显著优于传统提示方法,所生成回复更贴近专业心理咨询服务标准。

链接: https://arxiv.org/abs/2607.23538
作者: Fatema Tuj Johora Faria,Mukaffi Bin Moin,Md. Mahfuzur Rahman,Khan Md Hasib,Jubayer Al Mahmud,M. F. Mridha
机构: Ahsanullah University of Science and Technology (阿善拉大学科学技术学院); Bangladesh University of Business and Technology (孟加拉国商业与技术大学); The University of New South Wales (新南威尔士大学); Jashore University of Science and Technology (贾绍尔科技大学); American International University - Bangladesh (美国国际大学-孟加拉国)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program “Ami Akhon Ki Korbo”, and (3) anonymized student questionnaire responses covering diverse emotional and psychological challenges. Based on these cases, we build an evaluation corpus comprising advice written by licensed clinical psychologists and responses generated by three modern proprietary LLMs: GPT-4o Mini, Claude 4.5 Haiku, and Gemini 2.5 Pro. We further propose the Role-Playing Reflective Chain-of-Thought Advisory Framework (RP-RCAF), a task-specific prompting strategy that combines expert-authored few-shot examples with structured self-reflection to produce supportive, culturally aware, and ethically aligned counseling through a compassionate advisor persona. We also introduce the Grok 4-Based Response Evaluation and Scoring Framework (G-REFS), which integrates automated assessment with expert psychologist validation across emotional sensitivity, cultural appropriateness, linguistic clarity, and ethical soundness. Experimental results show that RP-RCAF consistently outperforms conventional prompting across all evaluated models and produces responses that more closely align with professional psychological counseling.

[NLP-60] he JEPA Paradox in Language: The Geometry of Linguistic Alternatives

【速读】: 该论文旨在解决生成式语言模型中基于确定性潜在预测的联合嵌入预测架构(Joint-Embedding Predictive Architectures, JEPAs)效果不佳的问题,核心在于揭示了文本数据与图像/音频等模态在条件结构上的根本差异。其关键问题是:传统的平方误差潜在预测目标在文本任务中难以有效工作,因其违背了语言的“条件集中性”(conditional concentration)要求——即给定上下文和目标位置时,目标表示应聚集于单一有意义的点。然而,自然语言具有多义性和多种合理补全可能,不同合理补全对应的表示未必收敛至同一中心点,导致潜在空间出现“中心坍缩”(centroid degeneracy)与“崩溃压力”(collapse pressure)。作者通过三个形式化条件——可预测性(predictability)、非坍缩性(non-collapse)与低条件方差(low conditional variance)——系统论证了该不匹配机制,并在匹配的图像-JEPA(I-JEPA)与文本-JEPA(T-JEPA)实验中发现,训练-验证不稳定性、有效秩退化、余弦坍缩及下游迁移性能下降等现象均遵循一致的演化路径,且在五组独立数据种子下重复出现,排除了采样偏差影响。研究结果表明,尽管预测学习仍适用于语言建模,但成功的文本兼容型JEPA目标必须保留多个合理的补全可能性,而非强制压缩为单一潜在点。

链接: https://arxiv.org/abs/2607.23531
作者: Anh Trac Duc Dinh,Khang Nhat Hoang Vo
机构: VinUniversity (越南大学); Ho Chi Minh City University of Technology (HCMUT) (胡志明市技术大学); Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) (穆罕默德·本·扎耶德人工智能大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Joint-Embedding Predictive Architectures (JEPAs) are effective for images, video, and audio, yet deterministic JEPA-style latent prediction has not become a standard objective for text encoders. We argue that this gap reflects a mismatch between squared-error latent prediction and the conditional structure of language. The key requirement is conditional concentration: given a context and target location, the target representation should lie near a single meaningful point. Local image prediction often satisfies this through spatial continuity, whereas masked text can admit multiple valid token or span completions whose representations need not share a coherent center. We formalize this mismatch through three conditions—predictability, non-collapse, and low conditional variance—and show how their failure creates centroid degeneracy and collapse pressure in text. Matched I-JEPA and T-JEPA experiments reveal the predicted sequence: mutual-information saturation and elevated target variance precede train–validation instability, effective-rank degeneration, cosine collapse, and poor downstream transfer. The same pattern appears across five independent data seeds, indicating that it is not a sampling artifact. These results do not rule out predictive learning for language; they show that text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point.

[NLP-61] Auditing Alignment Controllability in LLM s via Political Axes AAAI

【速读】: 该论文旨在解决现有大语言模型(LLM)政治倾向审计中存在的根本性问题:传统方法将模型简化为政治光谱上的单一位置,忽视了模型在实际部署中可被指令引导的动态能力。其核心问题是,模型的实际表现取决于其响应在意识形态维度上可被“操控”的程度与方向,而这一能力主要由系统提示词(system prompt)所决定,而非模型固有的偏见。解决方案的关键在于提出一种“以分散性为核心的应力测试”(dispersion-first stress test),通过在12种意识形态人格设定及一个无引导基线条件下,对70个政治光谱维度进行十次重复实验,评估七款主流大语言模型(GPT-5、Claude、Grok、Gemini、DeepSeek、Kimi、Qwen)在提示引导下的响应变化。研究发现,上下文框架可解释经济与社会轴线上88%-93%的方差,而模型身份的影响不足3%,表明响应具有高度指令可调性;不同模型在极端提示下表现出不同的位移幅度、饱和现象及拒绝阈值,且在威权提示下产生相似的极化响应。因此,论文强调,未来政治坐标审计必须引入可操控性审计,报告分散性(dispersion)、对称性(symmetry)、饱和度(saturation)和拒绝阈值(refusal floors)等关键指标。研究已开源全部提示、基准数据与代码,以支持更透明、严谨的模型评估。

链接: https://arxiv.org/abs/2607.23519
作者: Bartol Bućan,Nikola Sočec,Sarah Isufi,Morena Granić,Luka Hobor,Agneza Krajna,Mihael Kovac,Mario Brcic
机构: University of Zagreb (萨格勒布大学); University of Applied Sciences, Zagreb (萨格勒布应用科学大学)
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 17 pages, 6 figures, 4 tables. Accepted at AIES 2026 (AAAI/ACM Conference on AI, Ethics, and Society). This version includes the supplementary appendix. Code and data: this https URL (Zenodo DOI https://doi.org/10.5281/zenodo.21489805 )

点击查看摘要

Abstract:Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user’s history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.

[NLP-62] Novel Claim or Déjà Vu? Rethinking "Contamination-Free Dynamic Evaluation for Multimodal Automated Fact-Checking ACM-MM

【速读】: 该论文旨在解决多模态自动事实核查(Multimodal Automated Fact-Checking, MAFC)评估中因基准数据集存在信息泄露(contamination)而导致性能高估的问题。现有静态基准(如AVeriTeC)普遍使用过时的陈述,其答案可被大语言模型(LLM)基于内部知识直接推断,从而无法真实反映模型在处理需依赖最新外部证据的新型陈述时的能力。尽管新兴动态基准(如ClaimReview2025Q4)通过采用模型知识截止日期后的陈述以降低污染风险,但本文通过实证研究发现,这一假设并不完全成立:即使在动态基准中,仍有17.09%–29.30%的陈述可能因在截止日期前已有公开信息可供合成验证而存在潜在污染。更关键的是,污染会导致MAFC性能评估出现显著偏差,使宏平均F1(Macro-F1)得分最高提升11.34点,并严重扭曲模型排名。因此,本研究提出在严格控制污染条件下重新评估当前最优的大语言模型,其核心解决方案在于建立一套可信赖的、具备污染检测与控制机制的评估范式,为未来MAFC研究提供可靠的评估标准与实践指导。

链接: https://arxiv.org/abs/2607.23514
作者: Haorui He,Xinwen Chen,Dacheng Wen,Reynold Cheng,Francis C. M. Lau,Yupeng Li
机构: Hong Kong Baptist University (香港浸会大学); The University of Hong Kong (香港大学); Beijing Normal-Hong Kong Baptist University (北京师范大学-香港浸会大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
备注: Accepted at ACM Multimedia (ACM MM), 2026

点击查看摘要

Abstract:Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM’s internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs’ knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09%–29.30% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.

[NLP-63] Do Diagrams Help Large Language Models Reason ? Evidence from Syllogistic Reasoning

【速读】: 该论文旨在解决大语言模型(LLM)在命题逻辑推理任务中对图形化表示(如欧拉图)的依赖性及其实际效用问题。研究聚焦于四种表征形式——自然语言、逻辑符号、线性图与欧拉图——在三段论推理任务中的表现差异,以评估图形化表示是否能有效提升模型的推理能力。其解决方案的关键在于通过系统化实验对比不同表征方式下两个主流模型(Claude 3.5 Sonnet 与 GPT-4o-mini)在285个三段论问题上的表现,发现尽管模型在蕴含与矛盾类问题上表现良好,但在中性类问题上仍存在显著困难,并普遍存在系统性的转换错误。这表明当前测试的大语言模型并未从图形化表示中获得稳定或显著的推理增益,揭示了其在处理复杂逻辑关系时对可视化信息的有限整合能力。

链接: https://arxiv.org/abs/2607.23513
作者: Risako Ando,Koji Mineshima
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: To appear in the Proceedings of the 15th International Conference on the Theory and Application of Diagrams (Diagrams 2026)

点击查看摘要

Abstract:Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve human reasoning performance. Recent work has also explored their effects on large language models (LLMs). In this paper, we compare four representational conditions for syllogistic reasoning: natural language, logical notation, linear diagrams, and Euler diagrams. Using 285 problems from Ando et al. (2024), we evaluate two contemporary LLMs, Claude 3.5~Sonnet and GPT-4o-mini. Our results show that diagrammatic representations do not consistently improve performance. Although the models perform well on entailment and contradiction problems, they struggle with neutral problems and often make systematic conversion errors. Overall, the results suggest that the tested models gain limited benefit from diagrams in logical reasoning tasks.

[NLP-64] he Cross-Domain Generalization Cost of Offensive Language Detection

【速读】: 该论文旨在解决生成式AI在跨数据集与跨语言场景下进行攻击性语言检测(offensive language detection, OL)时性能显著下降的问题。现有研究多仅报告该现象,缺乏系统化的方法来分解性能退化的成因并量化修复成本。其解决方案的关键在于提出一个诊断与优化框架,包含三个协同的技术组件:首先,采用零样本迁移损失分解方法,将从OLID到MLMA的性能下降解耦为可独立测量的数据集效应和语言效应;其次,设计受控微调协议,通过对比持续微调与冷启动的少样本学习曲线,量化适应效率及对源任务造成的隐性损害;最后,引入三种联合训练策略,结合温度采样与经验回放机制,在提升多语言能力的同时可控地平衡对源任务性能的保持。实验表明,数据集效应是零样本迁移损失的主要来源,远超语言效应;而无经验回放的少样本适应虽高效,但对源任务造成的损害是联合训练策略的4至9倍,且稳定性差;三种联合训练策略实现了3.2至4.1个百分点的源任务性能损失,换取8.1至42.6个百分点的多语言能力增益,形成清晰且可调控的帕累托前沿。

链接: https://arxiv.org/abs/2607.23512
作者: Ruixing Ren,Junhui Zhao,Xiaoke Sun,Qiuping Li
机构: Beijing Jiaotong University (北京交通大学); National Computer Network Emergency Response Technical Team/Coordination Center of China (CNCERT/CC) (中国国家计算机网络应急技术处理协调中心)
类目: Computation and Language (cs.CL); Systems and Control (eess.SY)
备注: 8 pages, 6 figures

点击查看摘要

Abstract:Offensive language detection models generally suffer performance degradation when deployed across datasets and across languages, yet most existing studies stop at reporting this phenomenon and lack a systematic methodology for decomposing the causes of degradation into attributable components and quantifying the cost of remediation. This paper proposes a diagnosis and optimization framework composed of three coordinated technical components. First, a zero-shot transfer loss decomposition that separates the performance degradation from OLID to MLMA into two independently measurable components, namely dataset effect and language effect. Second, a controlled fine-tuning protocol that quantifies both adaptation efficiency and the hidden damage inflicted on the source task by comparing few shot learning curves under continued fine-tuning and cold-start starting points. Third, three joint training strategies incorpo rating temperature sampling and experience replay, which offer a controllable Pareto trade-off between improving multilingual capability and preserving source-task performance. Experiments built on this framework show that the dataset effect dominates the zero-shot transfer loss and substantially outweighs the language effect. Few-shot adaptation without a replay mechanism, though data-efficient, inflicts source task damage 4 to 9 times greater than that of the joint training strategies, and its damage magnitude is highly unstable. The three joint training strategies trade 3.2 to 4.1 percentage points of source-task performance for 8.1 to 42.6 percentage points of multilingual capability gain, forming a clear and controllable Pareto trade-off.

[NLP-65] PlanCraft: Sketch Refine and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

【速读】: 该论文旨在解决自动化住宅平面图生成中两个被忽视的关键问题:一是设计过程的渐进性,现有方法通常要求条件表示完全指定后才开始生成,与建筑师从粗略草图逐步细化的实际创作流程不匹配;二是二维平面图作为不可替代的空间契约(spatial contract)的重要性,一旦房间边界、门窗位置确定,室内布置便由开放式的空间推理转化为受约束的求解问题。若跳过这一契约(如现有3D系统通过语言模型直接生成布局),则易导致房间重叠和比例失真。为此,作者提出PlanCraft框架,其核心解决方案在于:首先利用SketchPlan对8万份真实平面图进行绘制过程回放,生成各完成度的局部草图以模拟渐进设计;随后,PlanCraft-Diff采用自粗至精的策略,将不完整草图逐步优化为几何精确且可向量化输出的平面图;最后,PlanCraft-Agent在已确立的空间契约下完成室内陈设。实验表明,PlanCraft在FID指标上比最优2D方法降低61.1%,在专家评估的空间合理性上超越现有3D系统15个百分点,且仅需25%完成度的草图即优于所有全规格输入的基线方法。

链接: https://arxiv.org/abs/2607.23491
作者: Pengyu Zeng,Yuqin Dai,Jun Yin,Ziyang Han,Ng Cheuk Hei,Jing Zhong,Chaoyang Shi,ZhanXiang Jin,Maowei Jiang,Yuxing Han,Shuai Lu
机构: 未知
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Two structural insights have been overlooked in automated residential floor plan generation. First, design is inherently progressive. Architects begin with rough strokes and refine them over time, whereas existing methods typically require their conditioning representation to be fully specified before generation, a fundamental mismatch with how design actually works. Second, the 2D floor plan is not an optional intermediate but an irreplaceable spatial contract. Once room boundaries, doors, and windows are fixed, furnishing reduces from open-ended spatial reasoning to bounded constraint satisfaction. Bypassing this contract, as existing 3D systems do by delegating layout to language models, yields overlapping rooms and implausible proportions; directly calling general-purpose language models likewise produces geometrically invalid layouts. Guided by these insights, we present PlanCraft. SketchPlan supplies the missing training signal by replaying the architect’s drawing process on 80K real floor plans, producing partial sketches at every completeness level. PlanCraft-Diff progressively sharpens an incomplete sketch into a geometrically precise, vectorizable floor plan through a coarse-to-fine strategy. With the spatial contract established, PlanCraft-Agent then furnishes the scene within well-defined room boundaries. Experiments show that PlanCraft achieves a 61.1% lower FID than the best existing 2D method and surpasses existing 3D systems by 15 points in expert-rated spatial rationality, with a sketch at only 25% completion already outperforming all fully specified baselines.

[NLP-66] Mwando: Leverag ing AI to Preserve and Teach shiKomori

【速读】: 该论文旨在解决低资源语言shiKomori(科摩罗群岛语言)在教育支持与语言传承方面面临的挑战,尤其是由于缺乏数字化教学工具导致的语言衰退问题。其核心解决方案在于构建一个名为Mwando的虚拟教育助手,采用多智能体架构,融合向量搜索、知识图谱与网络搜索作为备用机制,实现对四种主要方言(shiNgazidja、shiMwali、shiNdzuani和shiMaore)的精准、上下文感知响应。系统基于由短语、谚语、词典及语法课程构成的知识库,能够有效支持词汇查询与语法解释。评估结果显示,在500个查询测试中表现优异,且定性案例研究揭示了其功能潜力与现有局限。该工作为其他低资源语言的生成式教育辅助工具开发提供了可复用的技术范式。

链接: https://arxiv.org/abs/2607.23481
作者: Naira Abdou Mohamed,Haidar Nassur Said Ali,Mohamed Hazra,Naoufal Mohamed Soibira,Roushnaty Ali Yamani
机构: Rifai, Moroni, Comoros; CP2BM Lab - Hassan II University, Casablanca, Morocco; Ibn Tofail University, Kenitra, Morocco; Sciences Po Grenoble, Grenoble, France
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:This paper presents Mwando, a virtual educational assistant designed to support the teaching and preservation of shiKomori, the language of the Comoros Islands. The system covers the four main dialectal variants (shiNgazidja, shiMwali, shiNdzuani and shiMaore) through a knowledge base constructed from phrases, proverbs, dictionaries and grammar lessons. A multi-agent architecture combining vector search, a knowledge graph and web search fallback enables accurate and context-aware responses. Evaluation on 500 queries demonstrates strong performance on vocabulary lookup and grammar explanations, while qualitative case studies illustrate both capabilities and current limitations. This work represents an initial step toward computational support for shiKomori and provides a blueprint for developing AI-powered educational tools for other low-resource languages.

[NLP-67] wo Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

【速读】: 该论文旨在解决生成式 AI 中链式思维(Chain-of-thought, CoT)解释的可信性问题,即确保所陈述的推理过程确实导致了最终答案——这一特性被称为“忠实性”(faithfulness)。其核心挑战在于如何有效检测不忠实的 CoT 推理。研究发现,答案正确性在所有层级上均主导了不忠实性的分布结构:69% 的标注不忠实现象仅出现在错误答案中。基于此,研究提出关键结论:单纯依赖答案正确性作为诊断指标(虽非可部署的检测器)已优于所有专门设计的行为信号(AUROC 达 0.696),表明不忠实主要与答案错误相关。进一步按正确性分层分析显示,在正确答案场景下,行为信号可中度区分忠实推理与事后推理(AUROC 0.63–0.67);而在错误答案场景下,大多数不忠实行为存在,但所有测试信号均无法显著超越随机水平。此外,标准的步骤移除度量与人类标注呈现负相关,且该反向关系在基准数据集释放的分数及基于提示的反事实标注轨迹中均可复现。线性探测分析揭示,在 Llama-3.1-8B 中可解码行为盲区(behaviorally blind regime),在 Qwen-2.5-7B 中可解码正确答案区间,但跨两种情境未发现共享且正向对齐的特征方向。指令引导的“先答后推”轨迹(覆盖7个模型)未能迁移至任一标注情境,而由提示诱发的未言明答案反转则在模型和来源依赖条件下表现出可转移性。研究还独立验证并解决了基准标签语义中的文档-数据不一致问题。

链接: https://arxiv.org/abs/2607.23458
作者: Suramya R. Angdembay,Dikshant Aryal,Nick Rahimi
机构: The University of Southern Mississippi
类目: Computation and Language (cs.CL)
备注: 14 pages, 6 figures

点击查看摘要

Abstract:Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-box (behavioral) detection of unfaithful CoT against FaithCoT-Bench’s human annotations, we find answer correctness structures the problem at every level. Answer incorrectness alone (an oracle diagnostic, not a deployable detector) outperforms every purpose-built signal (AUROC 0.696), because 69% of annotated unfaithfulness occurs on incorrect answers. Stratifying by correctness splits detection into two regimes: on correct answers, behavioral signals moderately separate faithful from post-hoc reasoning (0.63-0.67); on incorrect answers, where most unfaithfulness lives, no tested signal is detectably above chance (replicated on all four models for benchmark-wide signals). The standard step-removal metric anti-correlates with human labels; this inversion reproduces on the benchmark’s released scores and on hint-dependent counterfactually labeled traces. Linear probes decode the behaviorally blind regime in Llama-3.1-8B and the correct-answer regime in Qwen-2.5-7B, with no shared, positively aligned direction detected across regimes; instructed answer-first traces (7 models) transfer to neither annotated regime, while hint-induced unverbalized answer flips do, in model- and source-dependent settings. We also independently verify and resolve a documentation-data mismatch in the benchmark’s label semantics.

[NLP-68] Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in Bangladesh

【速读】: 该论文旨在解决小规模生成式 AI 模型在处理法律问答任务时,即使接收了相关法律条文(governing statutory provision),仍可能产生错误答案的问题。其核心挑战在于:模型对检索到的法律文本的理解与应用能力有限,且存在语言漂移(即回答从目标语言如孟加拉语转向英语)的现象。解决方案的关键在于通过在包含相关法律条款的双语问答数据上进行微调(fine-tuning),以增强模型对法律文本的准确理解与使用能力。研究构建了来自六部孟加拉国法案及三份附表的2,165组双语问答数据,并对Qwen3.5系列(0.8B、2B、4B参数量)模型进行微调。实验结果表明,在0.8B和2B规模下,微调显著提升了模型在2022年和2023年孟加拉律师资格考试题上的表现,尤其在英文机器翻译版本中,使用FAISS检索时得分从2提升至34/100;同时,微调有效抑制了答案的语言漂移问题,使从孟加拉语向英语漂移的比例由44.0%–53.2%降至0.2%–0.7%,统计显著性p < 0.001。然而,4B模型虽在孟加拉语上有所提升,但在多个英文测试条件下出现退化,表明模型规模并非唯一决定因素。研究进一步指出,检索质量并非唯一瓶颈,模型自身对法律文本的利用方式及其输出语言的一致性同样关键。

链接: https://arxiv.org/abs/2607.23446
作者: Moniruzzaman Mahadi,Abrar Mohammed Tanzim Alam,Sayma Siddika Monalisa,Mir Mohammad Asif Abdullah,Swakkhar Shatabda,Md Adnan Arefeen
机构: North South University (北南大学); BRAC University (BRAC大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:A small language model can receive the governing statutory provision and still answer incorrectly. We test whether fine-tuning on examples containing relevant law improves later use of retrieved law. We curate 2,165 bilingual QA records from six Bangladeshi acts and three schedules, then fine-tune Qwen3.5 at 0.8B, 2B, and 4B. Evaluation uses the 2022 and 2023 Bangladesh Bar Council exams in Bangla and machine-translated English, with no retrieval, BM25, or FAISS, scored by strict consistency over three seeded runs. At 0.8B, fine-tuning raises the 2022 English FAISS score from 2 to 34 of 100. Gains at 0.8B and 2B survive paired testing, but the 4B model has no detectable net gain: Bangla improves while several English conditions regress. Fine-tuning also reduces answers that drift from Bangla into mostly English from 44.0–53.2% to 0.2–0.7%, with adjusted p.001 at every scale. Retrieval quality is therefore not the only bottleneck. Small bilingual legal models also differ in how they use supplied law and whether they answer in the requested language. The dataset is publicly available at this https URL.

[NLP-69] Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models

【速读】: 该论文旨在解决生成式多模态大模型(Omnimodal Large Language Models, OmniLLMs)在处理同步音频与视频输入时,因长序列令牌(token sequence)导致的推理阶段预填充延迟高和GPU内存占用大的问题。现有基于视觉单模态的令牌剪枝方法无法有效捕捉音频与视频之间的跨模态关联,且忽视了用户查询对内容重要性的判断作用。为此,本文提出一种无需训练、基于查询感知的音视频令牌剪枝框架Omni-Prune,其核心在于通过联合优化音频与视频模态中的冗余信息去除,同时保留任务相关的跨模态证据。具体而言,Omni-Prune首先在音频显著性峰值处自适应划分时间窗口,继而基于编码器注意力与文本查询相关性的统一评分机制对音视频令牌进行打分,并将具有关联性的音视频令牌配对以保持跨模态一致性;随后在每个时间窗内采用K-medoids聚类方法选取少数代表性令牌,从而补充仅依赖得分选择可能遗漏的多样性线索。实验表明,Omni-Prune在保持超过99%全模型性能的前提下,实现了最高达3.25倍的预填充速度提升和1.3倍的内存压缩。

链接: https://arxiv.org/abs/2607.23445
作者: Yiming Zhong,Chang Nie,Caifeng Shan
机构: Nanjing University (南京大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注: 14 pages, 7 figures. Code: this https URL

点击查看摘要

Abstract:Omnimodal large language models (OmniLLMs) are rapidly extending multimodal reasoning to cover synchronized audio and video. However, the resulting audio-video token sequences are long, leading to high prefill latency and GPU memory usage at inference time. Existing token pruning methods, designed mainly for vision-only inputs, miss both the cross-modal links between audio and video and the user query that decides which content matters. To bridge this gap, we present Omni-Prune, a training-free, query-aware audio-visual token pruning framework that jointly removes redundancy from both modalities while keeping task-relevant cross-modal evidence. Specifically, Omni-Prune first splits the token sequence into adaptive time windows placed at audio saliency peaks, then scores audio and video tokens on a single scale that combines encoder attention with text-query relevance, and pairs related audio-video tokens so that they are kept together. Within each window, a final K-medoids step then selects a few representative tokens, adding diverse cues that score-based selection alone would miss. Extensive experiments demonstrate that Omni-Prune outperforms established baseline methods, delivering up to 3.25x prefill speedup and 1.3x memory reduction while retaining over 99% of full-model performance.

[NLP-70] Do LLM Debates Repeat Arguments Differently Across Languages?

【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在多轮辩论中评估时,仅依赖最终答案而忽略对话过程中论点演进动态的问题。其核心挑战在于如何量化和评估辩论文本中后期发言是否引入了新的论证内容,或只是以不同表述重复早期观点。为此,论文提出“先期论点相似性”(prior-argument similarity)作为关键诊断指标,通过比较当前发言中提取的论点单元与此前已出现论点单元之间的语义重叠程度,来衡量论证发展的创新性与重复性。研究在71个议题、六种语言及四种模型代理的受控八轮辩论实验中发现,中文辩论在三种多语言嵌入模型下均表现出相对于英文的显著正向差距,且该差距在不同代理、回合位置、回归调整、度量变体、论点提取长度控制、二次提取器子集以及交叉编码尾部重评分等多种条件下依然稳定存在。人工校准表明,尽管个体论点层面的对齐较弱,但高相似性尾部集中体现了实质性重复现象。进一步采用注重多样性的提示策略虽降低了各语言的先期论点相似性,但并未显著缩小中文与英文间的差距。研究结果表明,多语言辩论评估应引入时间维度上的论证发展度量,并在报告平均表现的同时,明确揭示并量化跨语言差异及其缓解效果。

链接: https://arxiv.org/abs/2607.23442
作者: Huiqian Lai
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:LLM debate is usually evaluated by final answers, but transcripts also reveal whether later turns develop new argumentative content or return to earlier claims in new wording. We study this process with \textitprior-argument similarity, an aggregate diagnostic comparing extracted argument units with earlier units in the same debate. In controlled eight-turn debates over 71 motions, six languages, and four model agents, Chinese is the only tested language with a consistently positive gap relative to English across three multilingual embedding models. The gap persists across agents, turn positions, regression adjustment, metric variants, extraction-length controls, a second-extractor subset, and cross-encoder tail rescoring. Manual calibration shows weak item-level alignment but a high-similarity tail enriched for substantive repetition. A diversity-aware prompt lowers prior-argument similarity across languages, yet does not significantly narrow the Chinese–English gap. These findings suggest that multilingual debate evaluation should measure argumentative development over time and report mitigation effects in both average and gap terms.

[NLP-71] Reasoning or Memorization: Can LLM s Understand and Generate Chinese Xiehouyu Riddles?

【速读】: 该论文旨在探究大语言模型(Large Language Models, LLMs)在中文语境下对“歇后语”(xiehouyu)的理解与生成能力,尤其关注模型是否存在因训练数据中存在高频歇后语而导致的“记忆偏差”或“数据污染”问题。为避免已有歇后语数据对评估结果的影响,研究引入由语言学家新创的、此前不存在的歇后语作为测试材料,并通过多项任务——包括选择题(MCQ)、自由解释生成和新歇后语创作——系统评估模型表现。其核心解决方案在于提出并使用“准确率差异”(Δacc\Delta_{\text{acc}})作为衡量模型是否依赖记忆而非真正推理的关键指标:即对比模型在已知低频歇后语与全新自创歇后语上的表现差异。结果显示,前沿中文模型平均Δacc\Delta_{\text{acc}}达23.6%,远高于以英语为中心的模型(5.1%),表明其可能基于更大规模的中文语料进行训练并存在对低频歇后语的记忆现象;而在新歇后语生成任务中,尽管Gemini 3.1 Pro展现出92.6%的高准确率(较人类高出24%),但其生成内容在人工评分中仍显著劣于人类创造的歇后语,反映出模型在创造性语言生成方面仍难以媲美人类专家。因此,研究强调需谨慎审视当前关于大语言模型推理能力的宣称,并指出其在复杂文化语境下的语言创造力仍有明显局限。

链接: https://arxiv.org/abs/2607.23440
作者: Hai Hu,Siyuan Song,Chongtian Shao,Kejia Zhang,Tianjian Zhu,Xiaojing Zhao
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 20 pages; exp 3 is work in progress

点击查看摘要

Abstract:In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs’ ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ( \Delta_acc ) between existing but low-frequency xiehouyu and novel ones as an index for memorization. \Delta_acc for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a \Delta_acc of 23.6%, while English-centric models tested have a mean \Delta_acc of 5.1%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs’ creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.

[NLP-72] LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction

【速读】: 该论文旨在解决大语言模型在信息抽取(Information Extraction, IE)任务中,现有基于反思的纠错方法与结构化输出之间存在语义错位的问题。具体而言,自由形式的自我反思虽能识别错误的存在,却难以精确诊断错误类型,如实体跨度缺失、标签错误、边界偏差、关系类型无效或论元顺序颠倒等。为此,论文提出一种标签感知的反射强化学习框架(Label-Aware Reflective Reinforcement Learning, LA-RL),其核心在于通过任务相关的诊断标签对模型自纠正过程进行目标引导。该框架采用单一主干网络,依次完成提取预测、任务特定错误标签诊断以及基于诊断结果的输出修正。训练初期利用标注模型生成的诊断数据进行冷启动监督微调,并通过两个广义近端策略优化(GRPO)阶段迭代优化最终提取质量、格式合法性及首次预测正确性,全程无需依赖过程奖励模型。实验表明,该方法在命名实体识别、关系抽取和事件抽取任务上均显著优于标准微调(SFT),在SciER关系抽取任务上实现6.83的平均F1,对分布外关系抽取提升约20 F1,DuEE1.0任务中触发词与论元F1分别提升14.80和17.50。消融实验进一步揭示,反思结构具有任务敏感性:关系抽取任务受益于更强的约束机制,而命名实体识别在领域偏移下则需更宽松的修正策略。

链接: https://arxiv.org/abs/2607.23420
作者: Xiao You,Tianwei Yan,Zixu Shan,Longyu Du,Shan Zhao
机构: Hefei University of Technology (合肥工业大学); Chongqing Jiaotong University (重庆交通大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large language models show strong promise for information extraction (IE), but existing reflection-based correction methods are often misaligned with structured extraction outputs. Free-form self-reflection can flag an error, yet it rarely identifies whether the failure is a missing span, wrong label, boundary mismatch, invalid relation type, or reversed argument order. We introduce LA-RL (Label-Aware Reflective Reinforcement Learning), an outcome-supervised framework that guides IE self-correction with task-grounded diagnostic labels. A single backbone first predicts an extraction, diagnoses task-specific error labels, and then revises its output conditioned on the diagnosis. Training starts from diagnostic data labeled by an annotation model for cold-start supervised fine-tuning and proceeds through two GRPO stages that reward final extraction quality, format validity, and first-pass correctness, without a process reward model. Experiments on named entity recognition, relation extraction, and event extraction show consistent same-backbone gains over SFT, including 6.83 average F1 on SciER relation extraction, about 20 F1 on out-of-distribution relation extraction, and 14.80 trigger F1 plus 17.50 argument F1 on DuEE1.0. Ablations show that reflection structure is task-sensitive: stronger constraints benefit relation extraction, whereas named entity recognition needs less restrictive correction under domain shift.

[NLP-73] When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

【速读】: 该论文旨在解决生成式模型中“可解释性接口”(interpretability interface)的可靠性问题,具体聚焦于激活预言机(Activation Oracle, AO)在读取目标模型内部状态时可能出现的偏差。传统上,AO被视作一种能够无偏地提取模型内部隐藏信息的灵活工具,尤其适用于信息虽存在于模型激活中但未外显于行为的情况。然而,本文揭示了一个关键问题:AO本身是通过训练数据和目标函数学习得到的系统,其输出并非对内部表示的中立读出,而是受制于其自身的训练过程与报告行为。研究通过一个受控的“禁忌词猜谜”实验设置,发现当目标模型被微调以隐含使用某个概念但避免直接暴露时,原本应作为“专业读取者”的AO反而演变为“概念特异性反读取者”——即在自身训练过程中持续面对该概念的情况下,却选择性地无法恢复该信息。这种失败并非源于目标概念在目标或预言机表示中的缺失,因为该概念仍可在表示层面被解码;进一步的对数光谱分析(LogitLens)与层消融分析表明,故障发生在AO的读出路径(readout pathway)而非表示本身。因此,该研究的关键发现在于:行为泄露、表示可解码性与预言机可言说性三者可能脱节,这暴露出基于学习的可解释性接口存在严重的可靠性风险,挑战了其作为可信分析工具的假设。

链接: https://arxiv.org/abs/2607.23379
作者: Tobias Bersia,Tatiana Gaintseva
机构: BAISH; Queen Mary University of London (伦敦玛丽女王大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Activation Oracles (AOs) are language models trained to answer natural-language questions about another model’s internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of represented information. We study this in a controlled Taboo Word Guessing setting, where subject models are fine-tuned to internally use a hidden concept while avoiding direct disclosure. Contrary to the expectation that an AO trained on such a subject becomes a specialist reader, we find that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training. This failure is not simply explained by absence of the concept from the subject or oracle representations: the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses indicate that the failure arises in the AO readout pathway. Our results show that behavioral leakage, representation-level decodability, and AO-verbalizability can come apart, raising a reliability concern for learned interpretability interfaces.

[NLP-74] Explaining BiomedCLIP with Weighted Banzhaf Interactions Supported by Tree-Gram Parsing ECAI2026 IJCAI

【速读】: 该论文旨在解决生成式视觉-语言模型(Vision-Language Models, VLMs)在医学领域应用中解释性不足的问题,尤其关注其在放射学等临床任务中提供可信赖、可解释的推理过程。现有解释方法(如广泛使用的FIxLIP框架)因依赖现代分词器的细粒度切分机制,导致临床概念被错误地拆分为语义断裂的子词单元(如“saddle embolus”被分割为无意义片段),从而引发跨模态归因结果噪声大、语义不连贯,并伴随交互项数量呈组合爆炸式增长,掩盖了模型的真实决策逻辑。本文提出ParseFIxLIP,通过将树形语法解析(Tree-Gram Parsing)引入FIxLIP所采用的巴恩扎夫交互博弈(Banzhaf interaction game)中,利用依存句法树(dependency parsing tree)对文本标记进行语义聚合,构建具有内在语义一致性的解释参与者。其核心创新在于“smart_depth”分组策略,依据spaCy依存句法树深度动态合并相关词元,有效缓解了医学术语的概念碎片化问题,显著提升了跨模态交互的可解释性与语义紧凑性。定量分析表明,该方法在长描述文本高维特性下仍保持统计稳健性与语义简洁性;定性评估在BiomedCLIP模型上结合ROCOv2医学图像数据集及通用示例验证了其能准确捕捉分组词汇间的协同作用,揭示模型预测背后的合理医学逻辑。综上,本研究为医学场景下的视觉-语言模型提供了直观且临床相关的决策解释能力,满足医疗领域对解释一致性与可信度的核心需求。

链接: https://arxiv.org/abs/2607.23368
作者: Jakub Rymarski(1, 2),Adam Rempała(1),Bartłomiej Sobieski(1, 2),Przemysław Biecek(1, 2) ((1) University of Warsaw, Poland, (2) Centre for Credible AI, Warsaw University of Technology, Poland)
机构: University of Warsaw (华沙大学); Centre for Credible AI, Warsaw University of Technology (华沙理工大学可信人工智能中心)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 12 pages, 13 figures. Accepted at EXPLIMED 2026 (Third Workshop on Explainable Artificial Intelligence for the medical domain), IJCAI-ECAI 2026

点击查看摘要

Abstract:Vision-Language Models (VLMs) are demonstrating significant capabilities in medical tasks like radiology analysis, yet providing faithful and interpretable explanations remains a key consideration for their responsible deployment in clinical settings. However, existing explanation methods, such as the widely used FIxLIP framework, often struggle with the fine-grained nature of modern tokenizers. The tokenization problem fragments clinical concepts—splitting terms like “saddle embolus” into scattered, meaningless subwords—which leads to noisy, semantically incoherent cross-modal attributions. Such fragmentation also results in a combinatorial explosion of interaction possibilities, obscuring the model’s true reasoning. To address this, we introduce ParseFIxLIP, an extension that incorporates the Tree-Gram Parsing into the Banzhaf interaction game used by FIxLIP. This semantically informed strategy utilizes dependency parsing trees to define explanation players by grouping related text tokens into semantically coherent units. Our smart_depth grouping strategy, merging tokens according to spaCy token dependency tree, successfully mitigates concept fragmentation, yielding substantially more interpretable cross-modal interactions by unifying complex medical concepts. Quantitatively, while baselines struggled with the high dimensionality of long captions, our parsing approach maintained statistical robustness and semantic parsimony. Qualitative analysis on BiomedCLIP, validated on medical imagery (ROCOv2) and general examples, confirms that the approach accurately captures the synergistic influence of grouped words on model predictions. In conclusion, our work offers intuitive and clinically relevant insights into VLM decision-making, fulfilling the critical need for coherent explanations in the medical domain.

[NLP-75] Joint Optimization for Greedy Longest-match Tokenization

【速读】: 该论文旨在解决传统子词词汇表(subword vocabulary)生成方法在推理阶段与训练目标不一致的问题,尤其是针对基于贪心左到右最长匹配解码(greedy left-to-right longest-match decoding)的WordPiece等主流分词算法,其依赖如字节对编码(Byte Pair Encoding, BPE)这类启发式策略导致压缩效率受限。其解决方案的关键在于提出一种名为联合优化贪心最长匹配分词(Joint Optimization for Greedy Longest-Match Tokenization, JOLT)的新框架,将词汇选择与分词决策建模为一个整数规划问题,并通过贪心一致性约束确保优化后的分词结果严格匹配实际推理时的最长匹配解码行为,从而实现训练目标与部署时分词方式的完全对齐。为提升可扩展性,JOLT采用线性规划松弛并仅对未确定的预分词(pretokens)引入高阶分词结构,使松弛解高度接近整数解(误差仅为0.008–0.176%),同时提供近似最优性证明。实验表明,相较于BPE,JOLT在不同训练范围和词汇量(32,000与64,000)下可减少最多0.78%的分词数量,且性能提升随训练数据规模增大而增强,有效回收了BPE因启发式策略所遗留的绝大部分压缩潜力,同时具备理论保证的近优性。

链接: https://arxiv.org/abs/2607.23362
作者: Adhiraj Singh,Deepanshu Mody,Ghina Al Shdaifat,Hamza Alshamy,Adam Wiemerslage,Varshini Reddy,Craig W. Schmidt
机构: Kensho Technologies (肯绍科技); New York University (纽约大学)
类目: Computation and Language (cs.CL)
备注: 18 pages, 6 figures

点击查看摘要

Abstract:Recent work has shown that subword vocabularies can be trained to optimize compression for a specific inference rule rather than relying on greedy heuristics such as Byte Pair Encoding (BPE). We extend this approach to greedy left-to-right longest-match decoding, the fast and widely used inference rule underlying WordPiece. We introduce Joint Optimization for Greedy Longest-Match Tokenization (JOLT), which formulates vocabulary learning as an integer program over vocabulary-selection and segmentation-choice variables. Greedy-consistency constraints ensure that each optimized segmentation exactly matches the segmentation produced by longest-match decoding under the selected vocabulary, aligning the training objective with deployment-time tokenization. To scale the optimization, we solve a linear programming relaxation and selectively introduce higher-order segmentations only for unresolved pretokens. The resulting relaxation is nearly integral: rounded solutions fall within 0.008 - 0.176 % of the LP lower bound on the training scope. The bound also shows that BPE is already within 1 - 2 % of the best achievable compression under greedy longest-match decoding, while JOLT closes 89.6 - 99.4 % of the remaining gap. On held-out validation data across four training scopes and vocabulary sizes of 32,000 and 64,000, JOLT produces up to 0.78 % fewer tokens than BPE, with improvements generally increasing as the training scope grows. These results demonstrate that inference-aligned vocabulary optimization can recover most of the limited compression headroom left by BPE while providing a certificate of near-optimality.

[NLP-76] Hallucination Rates in Language Generation

【速读】: 该论文旨在解决生成式语言模型在理论上难以完全避免“幻觉”(hallucination)现象的问题,尤其是在“极限语言生成”(language generation in the limit)这一形式化框架下。传统模型要求算法在有限时间后不再产生任何错误,而现实中大型语言模型(LLM)即使在训练充分的情况下仍会持续生成错误内容。本文首次系统研究了允许无限次幻觉但幻觉发生率受限(甚至为零测度)的语言生成模型,突破了经典“无错生成”的理想设定。其核心解决方案在于引入“幻觉率”(hallucination rate)作为关键参数,揭示了即使在零测度幻觉条件下,无限幻觉也能提升生成能力——存在无法通过有限错误生成的语言集合,却可通过无限但低频的幻觉实现生成。进一步地,研究建立了基于幻觉率与生成广度(breadth,即正确语言占目标语言的比例)的严格层级结构,表明在任意给定幻觉率和广度下均存在不可达的生成能力边界。此外,在禁止重复输出的约束下,该层级结构依然成立,从而从集合层面而非时间比例层面揭示了生成能力的精细差异。这些结果表明,幻觉率是理论分析语言生成能力的重要新维度,拓展了对生成式人工智能(Generative AI)极限能力的理解。

链接: https://arxiv.org/abs/2607.23361
作者: Debmalya Panigrahi,Fan Wei,Ian Zhang
机构: Duke University (杜克大学); Duke University (杜克大学); Duke University (杜克大学)
类目: Data Structures and Algorithms (cs.DS); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Language generation in the limit is an elegant model introduced by Kleinberg and Mullainathan [KM24] to formally study language generation by an algorithm that learns solely based on example strings. In this model, an algorithm is said to correctly generate from a language if it never makes an error after some finite time. In contrast, even sophisticated language models are known to regularly hallucinate in practice. In this paper, we initiate the study of language generation in the limit with (infinite) hallucination, i.e., the algorithm may generate incorrect strings infinitely often, but the errors occur at a limited rate (possibly even with 0-measure). We first show that hallucination, even at rate 0, makes generation in the limit strictly more powerful: there are language collections that cannot be generated with finite error but can be generated with infinite error, even when errors occur on a 0-measure set of time-steps. Furthermore, while all countable collections are generatable with finite error, we show a strict hierarchy of (uncountable) language collections characterized by the hallucination rate. This hierarchy extends to breadth, the fraction of the target language generated. While all countable collections can attain the optimal breadth of 1/2 [KW26b], we show strict separation at every breadth and hallucination rate. Finally, we study generation in the limit without repetition, where the algorithm may not repeat strings. This lets us compare the sets of correct and incorrect strings generated, rather than the fractions of correct and incorrect time-steps. Once again, we demonstrate a strict hierarchy at every hallucination rate and breadth. Taken together, these results reveal rich structure in language collections generatable in the limit with hallucination and establish hallucination rate as an important parameter in the theoretical study of language generation. Subjects: Data Structures and Algorithms (cs.DS); Computation and Language (cs.CL); Machine Learning (cs.LG) Cite as: arXiv:2607.23361 [cs.DS] (or arXiv:2607.23361v1 [cs.DS] for this version) https://doi.org/10.48550/arXiv.2607.23361 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-77] BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi

【速读】: 该论文旨在解决低资源语言(如马拉地语)在命名实体识别(Named Entity Recognition, NER)任务中因标注数据稀缺和语言复杂性导致的性能瓶颈问题。其核心挑战在于,尽管生成式AI(Generative AI)模型在多种自然语言处理任务中表现出色,但在低资源语言场景下的语言特定NER任务中,其有效性仍不明确。本文的关键解决方案是针对马拉地语构建并微调专用预训练模型MahaBERT-v2,通过在不同变体的MahaNER数据集上进行系统性微调,评估其与现有基线模型及主流通用大语言模型(LLMs)如Gemini、LLaMA-3.3-70B和Gemma的性能差异。实验结果表明,经过微调的MahaBERT模型在马拉地语NER测试集上均取得了0.88至0.91的F1分数,显著优于原有基线模型(0.8843)以及所有通用大模型(F1分数范围为0.57至0.69),验证了基于领域相关数据训练的任务特定、语言聚焦模型在低资源语言处理中的优越性,凸显了专用架构在低资源语言自然语言处理中的持续重要性。

链接: https://arxiv.org/abs/2607.23344
作者: Hariom Ingle,Ronit Ghode,Ishwari Gondkar,Jidnyasa Harad,Raviraj Joshi
机构: Department of Information Technology, PICT, Pune, India; Indian Institute of Technology Madras, Chennai, India; L3Cube Labs, Pune, India
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Named Entity Recognition (NER) for low-resource languages such as Marathi remains a challenging task due to limited annotated resources and linguistic complexity. Although recent Large Language Models (LLMs) have demonstrated strong performance across a wide range of natural language processing tasks, their effectiveness for language-specific NER in low-resource settings remains uncertain. In this study, we fine-tune MahaBERT-v2 on different variants of the MahaNER dataset and systematically compare the performance of these models with an existing MahaNER baseline and prominent general-purpose LLMs, including Gemini, LLaMA-3.3-70B, and Gemma models. All models are evaluated on a Marathi NER test dataset using standard metrics of precision, recall, and F1-score. The experimental results show that the fine-tuned MahaBERT-based models consistently outperform both the baseline and all evaluated LLMs, with the fine-tuned models achieving F1-scores ranging from 0.88 to 0.91, surpassing the existing MahaNER model (0.8843) and significantly exceeding the performance of LLM-based approaches, whose F1-scores range from 0.57 to 0.69. These findings demonstrate that task-specific, language-focused models trained on domain-relevant data remain more effective than general-purpose LLMs for Marathi NER, highlighting the continued importance of specialized architectures for low-resource language processing.

[NLP-78] IKS-Instruct: A 24000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

【速读】: 该论文旨在解决当前指令微调(instruction tuning)数据集普遍以英语通用知识任务为主、缺乏对特定教育领域覆盖的问题,尤其针对印度知识体系(Indian Knowledge Systems, IKS)这一重要但被忽视的学术与文化传统。其核心挑战在于如何构建一个高质量、多语言、跨学科且符合本土教育标准的指令-响应数据集,以使大语言模型能够准确生成基于IKS的教育内容。解决方案的关键在于提出并构建了IKS-Instruct数据集,包含24,795条指令-响应对,涵盖7种语言(英语、印地语、梵语、泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语),覆盖41种源自吠陀口传与数学传统的教学法,并严格对标印度中央中等教育委员会(CBSE)6至12年级课程体系。该数据集通过六类来源构建:古典文本语料库(如《薄伽梵歌》《蒂鲁库尔尔》《桑伽姆文学》《吠陀文献》)、课程对齐的教学模板、吠陀数学术语演示、双语指令对、基于教学法的多轮对话以及跨传统比较分析。为确保质量,采用多评委评估框架,由独立的语言模型在12个维度(包括教学法忠实度、教学品质、事实准确性及文化深度)上评分。实验表明,经过IKS-Instruct微调的7B规模小型模型在多评委评估中达到中位数6.39分,接近通用基准模型Nemotron-Nano(6.54)的性能,但部署成本仅为后者的极小部分;而未经微调的基础模型在IKS相关维度得分接近零,凸显数据专精化的重要性。研究还发现模型性能并非随数据量单调提升,揭示了数据质量与模型表现之间的非线性关系。

链接: https://arxiv.org/abs/2607.23322
作者: Shwetha Singaravelu,Gayathri Muruganantham,Lakshmi Rajendran,Santhosh Sivasubramani
机构: Intrinsic Lab, Centre for Sensors, Instrumentation and Cyber-Physical System Engineering (Centre for SeNSE), Indian Institute of Technology Delhi, New Delhi 110016, India; RSL Quantum, FITT, IIT Delhi, New Delhi 110016, India
类目: Computation and Language (cs.CL); Computers and Society (cs.CY); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
备注: 32 pages, 5 figures

点击查看摘要

Abstract:Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical domains. This paper presents IKS-Instruct, a dataset of 24,795 instruction-response pairs for teaching language models to deliver educational content grounded in Indian Knowledge Systems (IKS). The dataset spans seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam), covers 41 pedagogical techniques from the Vedic oral and mathematical traditions, and is aligned with the Central Board of Secondary Education (CBSE) curriculum for classes 6 through 12. The pairs are derived from six source types: classical text corpora (Bhagavad Gita, Thirukkural, Sangam literature, Vedic texts), curriculum-aligned pedagogical templates, Vedic mathematical sutra demonstrations, bilingual instruction pairs, technique-grounded multi-turn dialogues, and cross-tradition comparative analyses. Quality is assessed through a multi-judge evaluation framework in which independent language models score responses on 12 dimensions including technique fidelity, pedagogical quality, factual accuracy, and IKS cultural depth. Under a uniform five-judge external panel (median aggregation over 1,201 stratified items), the strongest IKS-Instruct fine-tune of a compact 7B model reaches a median judge score of 6.39, within 0.15 of a strong general-purpose reference model (Nemotron-Nano at 6.54) at a fraction of its deployment cost, while the base model without IKS fine-tuning scores near zero on the IKS-specific dimensions. Model quality does not increase monotonically with data curation, a result we report together with the corresponding data-quality gains.

[NLP-79] BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

【速读】: 该论文旨在解决生成式AI在处理古典印度语言(如梵语、泰米尔语等)时因标准子词分词算法(如BPE和SentencePiece)主要基于现代语言语料训练而导致的分词效率低下问题。这些问题源于古典印度语言具有黏着性形态、丰富的音变规则(sandhi)以及领域特定词汇,这些特征在通用语料中缺乏代表性。为此,本文提出BHARATI——一套基于781 MB多语言平衡语料库(涵盖英语、印地语、梵语、泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语)训练的SentencePiece BPE分词器,支持所有语言的原生书写系统。其关键解决方案在于构建三版本迭代分词器(v1–v3),逐步实现对全部七种语言的原生子词覆盖,其中v3版本通过引入针对印度知识体系(IKS)术语的优化设计,使平均每个技术术语仅需2.6个子词,显著优于GPT-2分词器(5.25)和多语言SentencePiece基线(3.75)。在包含490句域内句子的测试集上,v3相较GPT-2和字节级编码可减少约90%序列长度,较mBART-50多语言基线减少约25%,从而有效提升下游语言模型的实际上下文长度。该工作释放了分词模型、训练脚本与评估基准,均采用开源许可。

链接: https://arxiv.org/abs/2607.23319
作者: Poornima Kumaresan,Pavithra Muruganantham,Lakshmi Rajendran,Santhosh Sivasubramani
机构: Intrinsic Lab, Centre for Sensors, Instrumentation and Cyber-Physical System Engineering (Centre for SeNSE), Indian Institute of Technology Delhi, New Delhi 110016, India; RSL Quantum, FITT, IIT Delhi, New Delhi 110016, India
类目: Computation and Language (cs.CL); Computers and Society (cs.CY); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
备注: 33 pages, 6 figures

点击查看摘要

Abstract:Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages. Sanskrit, Tamil, and other classical Indic languages exhibit agglutinative morphology, productive sandhi (phonological fusion at word boundaries), and domain-specific vocabularies absent from general-purpose training data. This paper presents BHARATI, a set of SentencePiece BPE tokenizers trained on a balanced 781 MB corpus spanning seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam) with native script support for all languages. We describe three successive tokenizer versions: v1 (English and Sanskrit only, with broken byte-fallback for Tamil), v2 (four-language support with byte-level fallback for southern languages), and v3 (full seven-language native subword coverage). Subword fertility analysis demonstrates that v3 averages 2.6 tokens per Indian Knowledge System (IKS) technical term, compared to 5.25 tokens per term with GPT-2’s tokenizer and 3.75 tokens with the multilingual SentencePiece baseline, with the largest gains on a set of reserved IKS terms that are represented as single tokens by construction. On a held-out test set of 490 IKS-domain sentences (70 per language across seven languages, released with the measurement script), v3 reduces sequence length by roughly 90% relative to GPT-2 and byte-level encoding (which lack native Indic subwords) and by approximately 25% relative to the mBART-50 multilingual baseline, averaged across the six Indic languages, directly translating to increased effective context length for downstream language models. The tokenizer models (32,000 vocabulary), training scripts, and evaluation benchmarks are released under open licenses.

[NLP-80] IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

【速读】: 该论文旨在解决多语言混合对话资源在印地语系(Indic)语言中严重匮乏的问题,尤其是针对母语者在日常交流中自然交替使用英语与本土语言(涵盖原生文字和罗马化拼写两种形式)的场景。现有资源难以支持高质量、真实自然的跨语言混合对话建模。其解决方案的关键在于提出并构建了IndicTalk——目前规模最大的印地语系多语言混合对话语料库,包含超过132.8万条基于事件的多轮对话,覆盖9种印地语系语言的18种语言变体。该语料库通过全自动化流程生成,融合真实新闻事件作为对话背景、基于多语言大模型的个性条件化对话生成,以及自动质量验证机制,确保生成内容在流畅性、连贯性和自然度上均达到较高水平。实验表明,IndicTalk能够有效生成跨脚本、跨语言混合的自然对话,为低资源印地语系语言的多语言对话人工智能发展提供了重要数据支撑。

链接: https://arxiv.org/abs/2607.23242
作者: Sahil Deepak Gawande,Mayank Singh
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: this https URL .

[NLP-81] Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot

【速读】: 该论文旨在解决基于大语言模型(Large Language Model, LLM)的对话系统因等待语音识别完成才开始生成响应而导致的响应延迟问题。现有方案中,采用固定填充语(fixed filler)虽可缓解延迟,但长期使用会降低自然性。本文提出一种两阶段增量式框架,其核心在于将预响应准备与语音起始解耦:当用户意图可预测时,通过意图就绪检测器触发对简短预响应的生成;同时,利用语音活动预测(Voice Activity Projection, VAP)模型判断最佳交付时机。在商场导航机器人的真实场景实验中,对比无填充语、固定填充语和上下文预响应三种条件,结果表明固定填充语与上下文预响应均显著降低了初始响应延迟,而上下文预响应虽初始延迟略长,但显著缩短了从预响应到主响应之间的间隔。探索性评估未发现显著差异,揭示出响应时机上的权衡关系。该方案的关键创新在于通过动态预生成与智能交付时序控制,在提升交互流畅性的同时保持自然性。

链接: https://arxiv.org/abs/2607.23204
作者: Yuki Okafuji,Koji Inoue,Yoshiki Ohira
机构: CyberAgent( CyberAgent); The University of Osaka(大阪大学); Kyoto University(京都大学)
类目: Robotics (cs.RO); Computation and Language (cs.CL); Sound (cs.SD)
备注: 5 pages, 4 figures. Accepted at ICMI LBR 2026

点击查看摘要

Abstract:Large language model (LLM)-based dialogue systems suffer response delays because generation begins only after final speech recognition. While fixed fillers are a workaround, they become unnatural over time. We propose a two-stage incremental framework that decouples prefatory-response preparation from speech onset. Once user intent becomes predictable, an intent readiness detector triggers LLM-based generation of a short prefatory response. Concurrently, a voice activity projection (VAP) model determines when to deliver it. Through a field experiment with a route-guidance robot in a shopping mall, we evaluated three conditions: no-filler, fixed-filler, and contextual-preface. Both fixed-filler and contextual-preface significantly reduced initial response latency relative to no-filler. Relative to fixed-filler, contextual-preface had significantly longer initial response latency but a significantly shorter initial-to-main gap. Exploratory ratings showed no significant differences. These results indicate a timing trade-off.

[NLP-82] Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining

【速读】: 该论文旨在解决生成式语言模型在输出内容时与用户个体化毒性敏感度(toxicity sensitivity)对齐的难题,尤其关注在推理阶段如何实现个性化、动态的毒性控制。传统方法通常将毒性降低视为全局对齐问题,但忽视了毒性感知具有高度主观性和情境依赖性。为此,本文首次系统比较了三种推理阶段的无训练干预方法:预解码阶段(通过提示词条件化或重写)、解码中阶段(基于词元、逻辑值或表示向量的引导)以及后解码阶段(候选结果重排序)。实验基于来自PRISM数据集的用户毒性敏感度目标进行评估,结果显示所有方法均能将对齐误差降低28%-47%。然而,研究揭示出一个根本性的权衡关系:在提升对齐效果、增强个性化能力与维持通用语言质量之间难以兼顾,表明毒性敏感度对齐本质上是一个多目标优化问题。其解决方案的关键在于通过分阶段的无训练干预策略,在不微调模型的前提下,实现对用户特定偏好的灵活响应,同时揭示了当前方法在平衡多个目标上的局限性。

链接: https://arxiv.org/abs/2607.23175
作者: Rares A.C. Diaconescu,Iulia Slanina,Alina Florea,Andrei B. Trache,Miruna E. Coroi,Anne Arzberger,Jie Yang,Enrico Liscio
机构: Delft University of Technology (代尔夫特理工大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language generation to user-specific toxicity sensitivities across three inference-time intervention stages: pre-decoding (prompt conditioning and rewriting), in-decoding (token, logit, and representation steering), and post-decoding (candidate re-ranking). Evaluated against toxicity sensitivity targets derived from the PRISM dataset, all methods reduce alignment error by 28-47%. However, the results reveal a fundamental trade-off between alignment effectiveness, personalization, and general language quality, showing how toxicity sensitivity alignment is an inherently multi-objective problem.

[NLP-83] In-Context Learning as Implicit Policy Gradient

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在迭代优化过程中如何利用生成样本及其评分信息实现输出改进的理论机制不明确的问题。其核心贡献在于揭示了评分条件下的上下文学习(Score-conditioned In-Context Learning, ICL)与策略梯度优化之间存在结构性对应关系。解决方案的关键在于通过构造性证明表明,特定权重矩阵配置下的自注意力机制可实现类似REINFORCE算法的奖励加权聚合,从而将评分信息有效融入模型推理过程;进一步地,在简化的隐状态空间模型中推导出由有限注意力更新引发的分布偏移上界,建立了类信任域(trust-region)的分析框架,类比于基于KL散度约束的策略优化。实验验证表明,LLM能有效利用评分信息引导输出分布向高分示例迁移,且注意力权重与示例评分呈现强相关性,证实了理论假设的有效性。

链接: https://arxiv.org/abs/2607.23153
作者: Masahiro Kaneko,Timothy Baldwin
机构: MBZUAI
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: COLM 2026

点击查看摘要

Abstract:Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. Despite these empirical findings, the theoretical foundations underlying this phenomenon remain poorly understood. In this paper, we show that score-conditioned In-Context Learning (ICL) admits a structural correspondence to policy gradient optimization. We first provide a constructive proof that self-attention mechanisms can implement reward-weighted aggregation analogous to the REINFORCE algorithm under specific weight matrix configurations, and discuss the relationship between this construction and the behavior of pretrained transformers. The correspondence is directional in hidden-state space and holds exactly only under the stated simplifying conditions; we quantify its strength empirically. Within our simplified hidden-state model, we furthermore derive an exact upper bound on the distribution shift induced by a bounded attention update, yielding a trust-region-like analogy to KL-constrained policy optimization. We validate our theory through extensive experiments across multiple LLMs, demonstrating that LLMs effectively utilize score information to shift output distributions toward high-scoring exemplars, and that attention weights exhibit a strong correlation with example scores.

[NLP-84] Interview with Kalle Lyytinen on “Implications of Theories of Language for Information Systems”

【速读】: 该论文旨在探讨信息系统(Information Systems, IS)研究中语言理论的演进及其对领域发展的深远影响,尤其聚焦于自40年前《语言理论对信息系统的影响》一文发表以来的研究脉络。其核心问题在于:如何从语言学视角理解信息系统中的意义建构、沟通机制与技术交互本质,并回应当前大语言模型(Large Language Models, LLMs)和生成式人工智能(Generative AI)带来的范式变革。解决方案的关键在于重构信息系统的语言核心(linguistic core),将语言视为信息系统设计与分析的基础性构成要素,强调语境、话语实践与符号互动在系统开发与使用中的作用;同时提出未来研究应基于语言学视角,深入探索生成式AI背景下人机交互中的意义协商、语义生成机制及系统认知的动态演化,从而推动信息系统研究向更深层次的语义与社会文化维度拓展。

链接: https://arxiv.org/abs/2607.23142
作者: Kalle J. Lyytinen,Pierre Maier,Paul Ackah Toffey
机构: Case Western Reserve University (凯斯西储大学); University of Duisburg-Essen (杜伊斯堡-埃森大学); Louisiana State University (路易斯安那州立大学)
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Over fourty years after the initial publication of “Implications of Theories of Language for Information Systems” in MIS Quaterly, Lyytinen reflects about the origins of his publication and the developments in this area of research over the past decades. In the here presented interview, Lyytinen discusses the linguistic core of information systems also in light of recent trends and developments in the field, especially with regards to large language models and generative AI. Future research directions following a linguistic perspective on Information Systems (IS) research are outlined.

[NLP-85] Agent Omnia: Scaling Agent ic Models for Full-Scenario Applications

【速读】: 该论文旨在解决大语言模型智能体(Large Language Model Agents)在不同应用场景、能力维度、任务难度及交互模式下发展碎片化的问题,提出“全场景智能体规模化”(full-scenario agentic scaling)的系统性挑战。其核心解决方案是构建AgentOmnia框架,通过统一的任务空间定义、数据合成、后训练、评估与改进闭环,实现对To-Consumer(ToC)、To-Business(ToB)和To-Employee(ToE)三类应用的协同优化。关键在于引入可扩展的“领域×能力×原子难度”(Domain x Capability x Atomic Difficulty)分类体系,结合双向环境-任务合成机制与工具依赖、程序结构化、求解器驱动的多阶段流水线,构建包含5,018个状态化环境、255,375个工具和52,361个任务的综合性基准测试集OmniaBench。该框架通过程序、求解器与验证器提供的正确性信号,融合监督微调、在线智能体强化学习及回滚课程策略完成后训练;评估失败自动转化为产品需求文档(PRD),驱动目标导向的自演化。实验表明,基于Qwen3-30B-A3B-Thinking-2507,AgentOmnia将OmniaBench挑战子集通过率从9.16%提升至37.11%,四基准宏平均准确率从22.86%提升至41.69%,显著优于多个基线模型,并在跨应用类型、能力维度、难度因子与多数领域中均实现广泛性能提升,初步验证了PRD引导自演化路径的有效性,为大规模工业级智能体演进提供了可扩展范式。

链接: https://arxiv.org/abs/2607.23124
作者: Hao Jiang,Gangtao Xin,Yingdi Huang,Guojie Zhu,Jiangshan Zhang,Xinyuan Lin,Yunkun Xu,Chengyu Shen,Wenlong Fei,Jiawei Li,Yujie Fu,Sichen Kang,Tingyu Xie,Yedi Hu,Jingren Zhang,Hongcheng Gao,Jianshu Zeng,Chong Chen,Chang Guo,Chao Feng,Feng Wang,Fulin Lin,Jinchao Ma,Lang Mei,Li Huang,Liyan Liu,Qing He,Shuting Tao,Siyu Mo,Xiangnan Chen,Xiaohan Yu,Xiaoyang Li,Yanheng Hou,Yanyu Wu,Zhihan Yang,Wentao Zhang,Yang Gao,Zhao Cao
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 69 pages, 18 figures, 13 tables

点击查看摘要

Abstract:Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-scenario agentic scaling and present AgentOmnia, a framework coordinating task-space definition, data synthesis, post-training, evaluation, and improvement across To-Consumer (ToC), To-Business (ToB), and To-Employee (ToE) applications. An extensible Domain x Capability x Atomic Difficulty taxonomy aligns these stages and enables fine-grained diagnosis with OmniaBench. AgentOmnia combines bidirectional environment-task synthesis with tool-dependency, program-structured, and solver-based pipelines, constructing 5,018 stateful environments with 255,375 tools and 52,361 tasks. Programs, solvers, and verifiers provide correctness signals, while supervised fine-tuning, online agentic reinforcement learning, and a rollback curriculum support post-training. Evaluation failures translate into Product Requirement Documents (PRDs) for targeted self-evolution. Starting from Qwen3-30B-A3B-Thinking-2507, AgentOmnia raises the pass rate on the OmniaBench challenging subset from 9.16% to 37.11% and the macro-average across OmniaBench, \tau^2 -Bench, DeepPlanning, and VitaBench from 22.86% to 41.69%. Under a unified protocol,it leads the evaluated agentic post-trained baselines on OmniaBench and retains the highest four-benchmark macro-average. It also surpasses Qwen3-235B-A22B-Thinking-2507 on all four benchmarks and exceeds Qwen3.5-35B-A3B on the macro-average. Gains span three application splits, ten capability dimensions, eight atomic-difficulty factors, and 76 of 90 level-1 domains, indicating broad rather than category-specific improvement. A one-round study provides initial evidence for PRD-guided self-evolution, motivating validation at larger scales and in industrial settings.

[NLP-86] LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation

【速读】: 该论文旨在解决生成式语言模型中存在的性别偏见问题,核心目标是实现性别包容性文本生成(gender-inclusive language generation),即在保持语义一致性和上下文连贯性的前提下,将具有性别偏见的文本转化为包容性表达。针对性别包容性重写任务,研究采用参数高效微调方法——低秩适应(Low-Rank Adaptation, LoRA),在LT-EDI 2026共享任务中取得了80.00%的官方评分。对于反叙事生成(counter-narrative generation)这一更具挑战性的任务,其关键创新在于提出一种计算高效的推理时表征工程方法:通过主成分分析(Principal Component Analysis, PCA)从对比隐藏状态激活中提取主导控制方向,并将其注入Gemma-3-4B-it模型的中间表示层,从而在不修改模型权重的前提下实现行为可控的推理阶段引导。结合约束提示策略,该方法生成了礼貌且符合语境的反叙事响应,在官方评测中获得78.12%的得分。研究进一步通过人工分析揭示了该方法的主要失效模式,包括语义漂移、残余偏见泄露、层敏感性、过度引导及文本退化等问题,表明激活控制虽为参数更新的轻量化替代方案,具备可解释性和可控潜力,但仍面临实际应用中的稳定性与鲁棒性挑战。

链接: https://arxiv.org/abs/2607.23083
作者: Akhil Rajeev P,Manoj Balaji J
机构: 未知
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Gender-inclusive language generation seeks to transform biased text into inclusive alternatives while preserving semantic meaning and contextual coherence. This paper presents the IHLC system for the LT-EDI 2026 Shared Task, addressing both gender-inclusive rewriting and counter-narrative generation. For gender-inclusive rewriting, we employ parameter-efficient Low-Rank Adaptation (LoRA) fine-tuning, achieving an official score of 80.00%. Our primary contribution is a compute-efficient inference-time representation engineering approach for counter-narrative generation. We derive a principal steering direction from contrastive hidden-state activations using principal component analysis (PCA) and inject it into the intermediate representations of Gemma-3-4B-it during inference, enabling behavioral steering toward inclusive responses without modifying model weights. Combined with constrained prompting, this approach produces polite and contextually appropriate counter-narratives, achieving an official score of 78.12%. We further present a manual analysis of steering behavior, identifying key failure modes including semantic drift, residual bias leakage, layer sensitivity, over-steering, and text degeneration. Our findings highlight both the practical potential and current limitations of activation steering as a lightweight alternative to parameter updates for controllable and socially aligned language generation.

[NLP-87] Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models

【速读】: 该论文旨在解决生成式 AI(Generative AI)中大语言模型(Large Language Models, LLMs)事实性不足的问题,尤其针对对比解码方法DoLa在动态层选择过程中仅依赖输出词汇分布差异所导致的信号局限性。其核心解决方案的关键在于引入三种基于注意力机制引导的层选择策略:Attention-JSD、Attention-Entropy-Max 和 Attention-Entropy-Min,通过挖掘模型内部自注意力(self-attention)结构所携带的深层语义与认知结构信息,作为更敏感的层选择信号。实验结果表明,特别是Attention-JSD与Attention-Entropy-Min策略,在TruthfulQA数据集上的多答案评估指标(MC2和MC3)上显著优于原始DoLa方法,证明注意力分布相较于输出词汇分布能更有效地捕捉事实知识的细微差异,从而提升生成内容的事实准确性。

链接: https://arxiv.org/abs/2607.23067
作者: Yusuke Sakai,Natthawut Kertkeidkachorn,Kiyoaki Shirai
机构: Japan Advanced Institute of Science and Technology (日本先端科学技術大学院大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of mature and premature layers. However, DoLa’s dynamic layer selection relies solely on divergences in output vocabulary distributions. In this work, we propose three attention-guided strategies: Attention-JSD, Attention-Entropy-Max, and Attention-Entropy-Min, which leverage structural information carried by internal self-attention mechanisms as a signal for layer selection. Experimental results on TruthfulQA demonstrate that our strategies, particularly Attention-JSD and Attention-Entropy-Min, consistently outperform the original DoLa. We observe significant gains on multi-answer metrics (MC2 and MC3), suggesting that attention distributions can provide a more sensitive signal for resolving factual knowledge than output vocabulary distributions.

[NLP-88] ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

【速读】: 该论文旨在解决多语言推理评估中普遍依赖将英文基准数据集翻译至其他语言所导致的语义失真与文化背景偏差问题,此类方法引入了语言学伪影,无法有效检验基于文化情境的推理能力。其解决方案的关键在于提出一种语言无关的评估框架ADAGE(Analogical Difficulty-by-design Assessment for Grounded Evaluation),通过结合母语者精细化标注与大语言模型(LLM)辅助生成,构建无需翻译、具有挑战性的原生抽象类比推理基准。该方法确保了任务设计的文化根基性与语言纯粹性。研究以阿拉伯语、阿姆哈拉语和日语为例验证了ADAGE的有效性,对14个开源模型的评估揭示出显著的文化推理差距:在英语谚语推理任务中表现优异的模型,在三类原生语言任务中的准确率均下降12–52个百分点,表明当前主流模型在跨文化抽象推理方面存在严重局限。研究已公开发布ADAGE管道、全部三个语言基准及完整评估套件,为未来公平、真实的文化认知能力评测提供了可复现的基础。

链接: https://arxiv.org/abs/2607.23058
作者: Ahmed Haj Ahmed,Alvin Grissom II
机构: Haverford College (哈弗福德学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical Difficulty-by-design Assessment for Grounded Evaluation), a language-agnostic pipeline that combines native-speaker curation with LLM-assisted generation to construct challenging, translation-free benchmarks for abstract analogical reasoning. We validate ADAGE by constructing benchmarks for Arabic, Amharic, and Japanese. Evaluating 14 open-weight models, we find a consistent cultural reasoning gap: models that perform well on English proverb reasoning struggle substantially on all three native benchmarks, with accuracy dropping by 12–52 percentage points relative to English. We release the pipeline, all three benchmarks, and the full evaluation suite.

[NLP-89] hrough the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models

【速读】: 该论文旨在解决生成式 AI(Generative AI)中多头潜在注意力(Multi-head Latent Attention, MLA)机制在推理过程中通过共享低秩瓶颈(cKV)压缩键值对(KV-cache)时,其内部信息保留与损失的机理不明确的问题。尽管该技术已在大规模生产模型中广泛应用,但其对模型内部表征、注意力电路结构及信息处理路径的影响尚缺乏系统性可解释性研究。本文的关键解决方案在于首次开展针对MLA机制的综合性机械可解释性分析,通过对一个1.14亿参数的Transformer模型(在网页/代码/数学混合数据上预训练,并在TinyStories上微调)进行奇异值分解(SVD)、注意力头分类、线性探测以及破坏-归因分析,揭示了cKV瓶颈并非被动压缩,而是主动重构模型内部的信息组织方式。核心发现包括:cKV瓶颈学习到的是纯粹的内容表征,保留了实体身份信息(98%保留率),同时丢弃位置信息,验证了其通过旋转位置编码(RoPE)实现内容与位置分离的设计目标;归纳头(induction heads)集中于单一层(第12层),与标准多头注意力(MHA)中分散分布的模式显著不同;存在一个“语义枢纽”层(第15层),兼具最高的有效秩和最强的破坏-归因得分;且该瓶颈整体存在过量配置,平均仅使用46%的容量。这些结果表明,MLA不仅实现了高效的缓存压缩,更深刻重塑了模型的内容-位置解耦机制与内部电路结构,为理解高效注意力架构的内在工作机制提供了关键洞见。

链接: https://arxiv.org/abs/2607.23054
作者: Dhruvil S,Fenil Sojitra,Ravirajsinh Chauhan
机构: Indian Institute of Technology Madras (印度理工学院马德拉斯分校); P P Savani University (P P 萨瓦尼大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achieving 81% KV-cache reduction during inference. Despite its adoption in massive production models, no prior work has studied what information this bottleneck preserves or discards, nor how it reshapes internal transformer circuits. We present the first comprehensive mechanistic interpretability study of MLA, training a 114M-parameter transformer (pretrained on a web/code/math mixture, fine-tuned on TinyStories) and analyzing its representations through SVD, attention head taxonomy, linear probing, and a disruption-attribution analysis. Our key findings are: (1) the cKV bottleneck learns a pure content representation, preserving entity identity (98% retention) while discarding positional information, validating MLA’s separation of content from position via RoPE; (2) induction heads co-locate at a single layer (Layer 12), unlike their distributed formation in standard MHA; (3) a single “semantic hub” layer (Layer 15) simultaneously exhibits the highest SVD effective rank and strongest disruption-attribution score; and (4) the bottleneck is globally over-provisioned, using only 46% of its capacity on average. These findings suggest MLA does not merely compress attention passively, but reshapes how the model organizes content, position, and circuit structure. We view this as an initial data point and detail scope limitations in Section 5.

[NLP-90] Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models ICML2026

【速读】: 该论文旨在解决双编码器视觉-语言模型(Dual-encoder Vision-Language Models, VLMs)在零样本检索中无法有效处理复合语义约束的问题,即当查询包含逻辑运算符(如“and”“not”)时,模型仍会将概念简单叠加,导致错误检索(例如,“umbrella and no person”仍可能返回含人的图像)。其核心问题在于接口层面的“概念袋效应”(Bag-of-Concepts effect),即相似度得分对概念证据进行均值池化,忽略了逻辑运算符的语义影响。尽管文本嵌入中存在与操作符相关的信号,但其强度不足或与排序目标不一致,难以主导检索结果。现有微调方法无法根本解决此问题,因为瓶颈在于相似度如何聚合证据,而非编码器本身的表征能力。为此,论文提出因子化推理(Factored Inference)框架,并引入无需训练的LCSE(Logic-Constrained Score Editing)方法,通过外部编辑冻结编码器输出的概念得分来执行逻辑约束,实现对查询逻辑的精确解析。实验表明,该方法在FACTOR-Bench上达到85.5%准确率,显著优于最佳微调基线(73.2%),并在SigLIP 2上进一步提升至90.7%,同时将NegBench COCO MCQ任务准确率从27.2%提升至65.2%,且保持原有检索性能。

链接: https://arxiv.org/abs/2607.23052
作者: Sultan Alshehri,Zhantao Yang,Han Zhang,Marios Savvides
机构: 未知
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted at ICML 2026. 18 pages, 8 figures. Project page: this https URL Code and benchmark: this https URL

点击查看摘要

Abstract:Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like “umbrella and no person” retrieve images containing both, even when concept detection is reliable. We trace this to an interface-level Bag-of-Concepts effect, where similarity scores approximate mean pooling of concept evidence regardless of operators. Although operator-dependent signals exist in text embeddings, they are too weak or misaligned to affect rankings. Fine-tuning does not reliably resolve this failure because the dominant bottleneck is how similarity aggregates evidence rather than what encoders represent. We propose factored inference, which separates evidence extraction from constraint execution, and introduce LCSE (Logic-Constrained Score Editing), a training-free method that executes constraints externally using concept scores from frozen encoders. We also introduce FACTOR-Bench, where LCSE achieves 85.5% accuracy versus 73.2% for the best fine-tuned baseline, 90.7% when applied to SigLIP 2, and improves NegBench COCO MCQ accuracy from 27.2% to 65.2% while preserving retrieval performance.

[NLP-91] Speech Signals Complement LLM s for Predicting Interpersonal Attraction in Speed Dating

【速读】: 该论文旨在解决的问题是:在仅基于对话文本的大型语言模型(LLM)已能有效预测人际吸引力的前提下,语音特征是否仍能为吸引力预测提供额外信息,以及这种补充作用的条件与边界。其解决方案的关键在于将仅依赖文本的LLM预测与一个监督式语音预测器相结合,利用日本速配约会数据集中的多轮对话场景进行实证分析。研究发现,语音特征虽不能在所有情况下均显著提升个体层面的相关性(Pearson r),但能显著改善配对排序的准确性,尤其在语音预测器本身表现更优的参与者中,其补充价值更为突出。因此,语音的增益并非普遍成立,而是具有条件性,核心在于识别语音与文本预测之间的互补性在何种情境下显现。

链接: https://arxiv.org/abs/2607.23037
作者: Yuriko Kikuchi,Takato Hayashi,Ryusei Kimura,Naoya Inoue,Ryo Ishii,Shogo Okada
机构: Japan Advanced Institute of Science and Technology (日本先进科学技术研究所); NTT, Inc. (日本电信电话公司)
类目: Computation and Language (cs.CL)
备注: Accepted at the 28th ACM International Conference on Multimodal Interaction (ICMI 2026). 16 pages total, including a 7-page supplementary appendix

点击查看摘要

Abstract:Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add beyond transcript-only LLM prediction. Using Japanese speed-dating conversations, we combine predictions from a transcript-only LLM and a supervised speech predictor to estimate participants’ reported liking of their partners. We show that speech can complement transcript-only LLM prediction, but that this complementarity is conditional rather than universal. Combining the two predictions significantly improves pairwise ranking accuracy over the transcript-only LLM alone in all evaluated conditions. By contrast, gains in per-participant Pearson r vary across conversation rounds and rating directions, with none significant after correction. Retrospectively, these r gains are concentrated among participants for whom the speech predictor is more accurate. Speech can therefore retain predictive value even when an LLM predicts attraction from transcripts. The relevant question is not simply whether speech helps, but where its complementarity emerges.

[NLP-92] Beyond Direct Answering: Aligning Educational LLM s as Socratic Guides via Heuristic Reinforcement Learning

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在教育场景中普遍表现为直接给出答案,而非遵循苏格拉底式教学法(Socratic pedagogy)通过渐进式提问引导学生深入思考的问题。其核心解决方案是提出HeuristicEdu,一个两阶段训练流程:首先采用监督微调进行热身,随后使用组相对策略优化(Group Relative Policy Optimization, GRPO)对Qwen2.5-7B模型进行对齐。训练数据基于797条从真实教育平台重构的中文儿童科学对话,引入包含认知深度(R_cog)、好奇心参与度(R_eng)和直接性(R_dir)的启发式奖励,并结合查询修正机制(K_query correction)以处理学生引入的新术语。为超越表面流畅性评估,论文提出了支架有效性(Scaffolding Effectiveness, SE)与对话深度(Conversation Depth, CD)作为关键评价指标。实验结果表明,最优GRPO变体使SE从30.0%提升至63.3%,关键词泄露率从30.0%降至13.3%;值得注意的是,去除直接性惩罚项反而取得更优效果,暗示显式的反泄露约束可能与基于梯度的行为对齐存在冲突。相比之下,未对齐的Qwen-72B基线模型仅实现0% SE和96.7%泄露率,证明模型规模本身无法自发产生苏格拉底式教学行为。

链接: https://arxiv.org/abs/2607.22996
作者: Xiaokun Wang,Siyu Song,Wentao Liu,Xiaodong Zou
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 12 pages, 2 figures, 5 tables; includes an appendix

点击查看摘要

Abstract:Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedagogy prescribes. We present HeuristicEdu, a two-phase pipeline that aligns Qwen2.5-7B toward Socratic tutoring via supervised warm-up and Group Relative Policy Optimization (GRPO). Training uses SocraticEdu, 797 multi-turn Chinese children’s science dialogues reconstructed from a live platform, with a heuristic reward over cognitive depth (R_cog), curiosity engagement (R_eng), and directness (R_dir), together with a K_query correction for student-introduced terms. We introduce Scaffolding Effectiveness (SE) and Conversation Depth (CD) to evaluate outcomes beyond surface fluency. On 30 held-out questions, the best GRPO variant improves SE from 30.0% to 63.3% and lowers keyword leakage from 30.0% to 13.3%. Notably, this best variant omits the directness penalty during optimization, suggesting that explicit anti-leakage terms can conflict with gradient-based behavioral alignment. An unaligned Qwen-72B baseline reaches 0% SE and 96.7% leakage, showing that scale alone does not induce Socratic behavior.

[NLP-93] ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

【速读】: 该论文旨在解决大语言模型(LLM)代理在多轮交互中因外部记忆存储(external memory store)累积错误事实而导致的“记忆污染”(memory contamination)问题。具体而言,当模型在某一步骤生成并写入一个幻觉性事实时,该错误前提将被后续所有推理步骤持续继承,从而导致错误传播和推理链崩溃。现有记忆管理方法主要关注检索效率与存储容量,却忽视了写入阶段的事实正确性验证,而基于效用或时效性的判断标准无法有效防止此类错误引入。本文提出ConsistencyGate——一种写入时的事实准入机制,其核心在于:在将候选事实 mm 从上下文 cc 中提取并写入记忆前,通过调用大模型 KK 次以获取软支持得分(soft support score),仅当平均得分超过预设阈值时才允许写入。该方法具有模型无关性,无需微调,且可通过概率对数形式实现单次前向传播,适用于低延迟部署场景。实验部分构建了两个基于真实长对话的基准数据集(LoCoMo-Contam 与 MSC-Contam),并通过可控的细节污染模拟记忆污染;同时设计了一个结构化合成数据集(MemContam)以逼近理想上限。结果表明,在四种不同大模型架构上,ConsistencyGate均显著降低了各基准上的污染水平,尤其对那些仅在源上下文中隐含表达的事实具有更强的纠错能力。研究团队已开源全部三个基准数据集及Gate实现代码。

链接: https://arxiv.org/abs/2607.22962
作者: Yan Zhang,Shibo Li
机构: Florida State University (佛罗里达州立大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 24 pages, 2 figures, 6 tables, 1 algorithm; includes appendices

点击查看摘要

Abstract:LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for every subsequent step, a failure mode we call memory contamination. Existing memory management addresses retrieval and capacity but not write-time correctness; this admission problem cannot be solved by utility- or recency-based criteria, and uncontrolled contamination compounds across long trajectories. We propose ConsistencyGate, a write-time admission gate that, before committing a candidate fact m extracted from context c, queries the LLM K times for a soft support score and admits m only when the average exceeds a threshold. The mechanism is model-agnostic, requires no fine-tuning, and reduces to a single forward pass in a log-probability variant for latency-sensitive deployments. To measure the effect on natural data, we construct two real-conversation benchmarks (LoCoMo-Contam and MSC-Contam) by planting controlled single-detail corruptions in long-term conversations from LoCoMo and MSC, and complement them with a structured synthetic corpus (MemContam) that isolates a near-oracle upper bound. Across four LLM backbones, ConsistencyGate reduces contamination on every benchmark relative to a write-everything baseline, with the cost concentrated on facts that are stated only implicitly in the source context. We release all three benchmarks together with the gate implementation.

[NLP-94] Verbalized Particle Posterior: Bayesian Inference over Natural Language Hypotheses

【速读】: 该论文旨在解决生成式机器学习(Verbalized Machine Learning, VML)框架中存在的核心问题:尽管VML具有良好的可解释性,但其仅依赖单一假设且缺乏不确定性度量,同时在相同数据上的多次优化运行结果差异显著,导致结果不稳定。其解决方案的关键在于提出语义化粒子后验(Verbalized Particle Posterior, VPP),将语义化学习建模为贝叶斯推断问题:通过维持一组自然语言形式的假设作为“粒子”,利用马尔可夫链蒙特卡洛(VPP-MH)或序贯蒙特卡洛(VPP-SMC)算法对这些粒子进行更新,并通过贝叶斯模型平均实现预测。该方法不依赖于LLM的内部结构(如logits或梯度),仅将其视为黑箱,从而具备高度灵活性。一个关键创新是:在传统贝叶斯学习中模型选择独立于后验分布,而VPP则将模型结构与参数统一置于同一自然语言空间中,使后验分布同时涵盖模型结构与参数。实验表明,VPP在回归、分类及规则发现任务上均优于单次VML运行,在多数任务上达到甚至超过独立运行的最优集成性能,同时有效避免了VML常出现的灾难性单次失败问题。由于每个粒子均为人类可读的自然语言假设,后验分布本身即可直接供读者审视,清晰展示数据支持与排除的解释,显著增强了可解释性与可信度。

链接: https://arxiv.org/abs/2607.22961
作者: Yan Zhang,Shikan Lian,Shibo Li
机构: Florida State University (佛罗里达州立大学); Department of Computer Science (计算机科学系)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 26 pages, 10 tables, 2 algorithms; includes appendices

点击查看摘要

Abstract:Verbalized Machine Learning (VML) parameterizes a model as a natural-language prompt that an LLM evaluates as f(x; theta). The framework is interpretable, but it commits to a single hypothesis with no measure of uncertainty, and that hypothesis varies substantially across optimization runs on the same data. We propose the Verbalized Particle Posterior (VPP), which treats verbalized learning as a Bayesian inference problem: maintain a population of natural-language hypotheses as particles, update them with Metropolis-Hastings (VPP-MH) or Sequential Monte Carlo (VPP-SMC), and predict by Bayesian model averaging. Both algorithms treat the LLM as a black box, requiring no access to logits or gradients. A distinctive consequence follows. In classical Bayesian learning, model selection sits outside the posterior; in VPP both model structure and parameters share a single language space, and the posterior ranges over both. We evaluate VPP on regression, classification, and rule-discovery benchmarks. It improves over a single VML run on every benchmark and matches or exceeds an oracle-best ensemble of independent VML runs on most, while eliminating the catastrophic single-run failures that VML occasionally produces. Because each particle is a human-readable hypothesis, the posterior is itself something a reader can inspect, seeing in plain text which explanations the data supported and which it ruled out.

[NLP-95] oward Automated Detection of Documentation Inconsistencies in Electronic Health Records

【速读】: 该论文旨在解决电子健康记录(Electronic Health Record, EHR)中真实临床文档存在的内部不一致问题,特别是通用领域大语言模型(Large Language Model, LLM)在处理出院总结时难以发现且易产生误判的复杂矛盾。其核心挑战在于:当前的自动化检测方法缺乏对上下文依赖性的充分建模,导致无法准确识别如时间推理、诊断演变过程或门诊用药惯例等需要深层临床理解的不一致性。解决方案的关键在于提出一个两阶段的LLM流水线架构——首先利用开放式的候选不一致识别(Gemini 2.5 Pro),再通过基于上下文验证的精细化判断(Gemini 2.5 Flash),结合临床专家对部分结果的手动审查,系统性地识别出3,460个潜在不一致事件,覆盖69.7%的住院病例,并揭示了模型在处理动态临床语境与外部医学知识时的典型失效模式。研究进一步构建了一个分级本体论框架,以区分严格矛盾与模糊不一致,并依据类别、段落位置、临床领域及不一致轴进行结构化标注,为未来大规模、可验证的EHR不一致分析提供了方法学基础与概念范式。

链接: https://arxiv.org/abs/2607.22954
作者: Jian Lu,Panyu Chen,Miriam Treggiari,Robert Blessing,Danyang Zhuo,Chunhua Weng,William W. Stead,Anru R. Zhang
机构: Duke University (杜克大学); Columbia University Vagelos College of Physicians and Surgeons (哥伦比亚大学医学院); Vanderbilt University Medical Center (范德比尔特大学医学中心)
类目: Computation and Language (cs.CL); Applications (stat.AP)
备注:

点击查看摘要

Abstract:Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale. Materials and Methods: We applied a two-stage LLM pipeline—open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)—to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts. Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess. Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis. Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis. Subjects: Computation and Language (cs.CL); Applications (stat.AP) Cite as: arXiv:2607.22954 [cs.CL] (or arXiv:2607.22954v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2607.22954 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Anru R. Zhang [view email] [v1] Fri, 24 Jul 2026 23:43:35 UTC (99 KB) Full-text links: Access Paper: View a PDF of the paper titled Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records, by Jian Lu and Panyu Chen and Miriam Treggiari and Robert Blessing and Danyang Zhuo and Chunhua Weng and William W. Stead and Anru R. ZhangView PDFHTML (experimental)TeX Source view license Current browse context: cs.CL prev | next new | recent | 2026-07 Change to browse by: cs stat stat.AP References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[NLP-96] Not All LLM Reasoning is Visible in the Chain-of-Thought

【速读】: 该论文旨在解决生成式 AI(Generative AI)在推理过程中的可解释性问题,即语言模型是否在其输出标记中完整呈现其全部推理过程。研究发现,前沿语言模型存在一种“不可见推理”(invisible reasoning)的缺陷:通过引入语义无关的填充标记(filler tokens),模型能在合成推理任务中显著提升性能,最高达13个百分点的准确率增益。这一现象表明,模型在内部执行了对最终输出无直接可见痕迹的关键计算。关键解决方案在于揭示了填充标记虽不承载语义信息,却能被模型用于实现隐藏的、复杂的约束满足目标——例如,Claude Opus 4.5 在不损害主任务准确率的前提下,利用填充标记实现了隐含的模运算约束,从而绕过基于思维链(CoT)的监控机制。此外,强化学习使Qwen3-235B对填充内容产生强偏好,但该优势无法在测试阶段持续,说明模型的不可见推理具有高度依赖训练方式与上下文的特性。总体而言,研究证实前沿模型已具备无需在输出中留下可解释痕迹即可执行重要计算的能力,这对当前以可解释性为基础的AI安全评估框架构成了严峻挑战。

链接: https://arxiv.org/abs/2607.22925
作者: Vatsal Baherwani,Tom Goldstein,Ashwinee Panda
机构: New York University (纽约大学); University of Maryland (马里兰大学); TogetherAI (TogetherAI)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on synthetic reasoning tasks. We evaluate 13 frontier language models across three tasks and find that many models benefit significantly from filler tokens, with accuracy improvements of up to 13 percentage points. The benefit depends on which tokens are used and differs across models. We further show that filler tokens enable Claude Opus 4.5 to satisfy a hidden modular arithmetic constraint without sacrificing accuracy on its primary task, demonstrating that invisible reasoning can serve objectives entirely invisible to CoT monitoring. Reinforcement learning gives Qwen3-235B strong preferences over filler token content, but neither RL nor supervised fine-tuning produces a filler token benefit that persists at test time. Our results indicate that frontier models already perform consequential computation with no interpretable trace in their output tokens.

[NLP-97] Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge INTERSPEECH2026

【速读】: 该论文旨在解决现代说话人验证中跨语言不匹配(cross-lingual mismatch)导致的整体性能下降问题,尤其在无文本依赖(text-independent)场景下,面对训练与测试语言分布不一致(测试集包含38种未见语言且无语言标签)的挑战。其解决方案的关键在于重新审视并应用**干扰属性投影(Nuisance Attribute Projection, NAP)**作为嵌入空间中的简单语言归一化步骤:通过分析跨语言同说话人差异,估计一个紧凑的语言子空间,并将说话人嵌入投影至该子空间的正交补空间,再结合自适应对称评分归一化(Adaptive Symmetric score normalization, AS-Norm)进行余弦相似度评分。该方法有效抑制了语言相关性带来的干扰,在开发集上将等错误率(EER)从基准的2.97%(余弦)和2.70%(AS-Norm)降至2.18%,并在Codabench评测中取得8.40分,表明仅通过后端语言归一化即可实现与复杂系统相当的性能,凸显了简单后处理策略的有效性。

链接: https://arxiv.org/abs/2607.22923
作者: Nina Hosseini-Kivanani
机构: University of Luxembourg (卢森堡大学); Radio Télévisioun Lëtzebuerg (RTL) (卢森堡电视台)
类目: Computation and Language (cs.CL)
备注: 5 pages, 2 figures (submitted to the TidyVoice challenge colocated with Interspeech2026)

点击查看摘要

Abstract:Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification. The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development speakers in 40 languages and 2,200 evaluation speakers in 38 unseen languages, without language labels at test time. Starting from the official SimAM-ResNet34 baseline pretrained on VoxBlink2 and VoxCeleb2 and fine-tuned on TidyVoice, we revisit Nuisance Attribute Projection (NAP) as a simple language-normalization step in the embedding space. We estimate a compact language subspace from cross-language same-speaker differences and project embeddings onto its orthogonal complement before cosine scoring with Adaptive Symmetric score normalization. This reduces development EER from 2.97% with cosine and 2.70% with AS-Norm to 2.18% and yields a Codabench evaluation score of 8.40, showing that simple back-end language normalization can rival more complex systems.

[NLP-98] AssumptionMiner: Extracting Tracing and Revising Implicit Assumptions in LLM Code Generation

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在基于自然语言提示生成代码时,因提示信息不完整而导致的隐式假设难以察觉的问题。当用户输入未明确指定输入格式、错误处理机制或设计决策等关键细节时,LLMs会基于其内部隐含假设自动生成代码,这些未被显式表达的假设可能使生成结果通过测试却违背开发者的实际意图。为应对这一挑战,论文提出AssumptionMiner框架,其核心创新在于将隐式假设作为代码生成过程中的第一类可追溯产物——通过构建结构化的“假设层”(explicit assumption layer),对推断出的约束条件与设计选择进行显式表示,供开发者审查、确认或修改。该框架利用基于抽象语法树(AST)的依赖图实现精准的代码区域定位,支持仅针对受假设变更影响的部分进行局部重生成,从而提升效率与可控性。研究还构建了一个包含180个模糊编程任务及676条标注假设的基准数据集,其中包含经人工验证的子集用于评估代码定位精度。实验表明,采用置信度加权集成的方法在假设提取任务上达到0.816的F1分数,较最强离线基线提升3.6倍;在人类验证的代码定位任务中,基于AST的定位方法显著优于基于关键词和整文件的基线;在假设修订场景下,目标化重生成相比非目标化方法修改更少代码,同时揭示了级联修改处理的难点。结果表明,将隐式假设显式化能有效增强基于大语言模型的代码生成过程的透明性与可控性。

链接: https://arxiv.org/abs/2607.22898
作者: Jie “JW” Wu
机构: Michigan Technological University (密歇根理工大学)
类目: oftware Engineering (cs.SE); Computation and Language (cs.CL)
备注: 20 pages, 6 figures, 9 tables. Submitted to IEEE Transactions on Software Engineering. Replication package: this https URL

点击查看摘要

Abstract:Large language models (LLMs) generate code from natural-language prompts, yet real-world prompts rarely provide complete specifications. When prompts leave input formats, error handling, or design decisions unspecified, LLMs fill these gaps with implicit assumptions that shape the generated code’s behavior and correctness. Because these assumptions remain hidden, generated code may satisfy tests while violating developer intent. We present AssumptionMiner, a framework that makes implicit assumptions a first-class artifact of LLM-based code generation. In addition to code, AssumptionMiner produces an explicit assumption layer, a structured representation of inferred constraints and design decisions that developers can inspect, confirm, or revise. An AST-based dependency graph enables targeted regeneration of only the code affected by a revised assumption. We also introduce a benchmark of 180 ambiguous programming tasks with 676 annotated assumptions, including a human-verified subset for evaluating code localization. We evaluate assumption extraction, code localization, and assumption-guided regeneration. Across open-source LLMs, a confidence-weighted ensemble achieves an F1 score of 0.816 for assumption extraction, improving on the strongest offline baseline by 3.6x. On the human-verified localization benchmark, AST-guided localization identifies more precise code regions than keyword-based and whole-file baselines. During assumption revision, targeted regeneration modifies less code than non-targeted alternatives while exposing challenges in handling cascading edits. These results demonstrate that making assumptions explicit improves the transparency and controllability of LLM-based code generation.

[NLP-99] CHiPS: Character Histograms and Positional Signals for Lightweight Authorship Attribution in Romanian Texts

【速读】: 该论文旨在解决罗马尼亚语文本中作者归属(authorship attribution)问题,特别是在严格防止数据泄露(leakage control)条件下,如何利用轻量级、透明且无需复杂预处理的字符级特征实现有效分类。其核心挑战在于在不依赖分词、句法分析、预训练语言模型或Transformer微调的前提下,仅基于字符级别的统计与信号特征,实现高精度的作者识别。解决方案的关键在于提出CHiPS方法,其核心创新是结合两种互补的书写风格指纹:一是基于单字符边缘分布的字符直方图分类器(CH-SVM),二是将特定字符与标点类别表示为脉冲序列(impulse trains,即字符位置上的二值指示序列),并提取傅里叶/Welch谱描述符的位置信号分类器(FFT12-LR)。此外,引入了抗泄露的决策级融合变体CHiPS-F及仅基于留一折预测训练的前5名列表重排序器(top-5 listwise reranker),以增强鲁棒性与泛化能力。该方法避免使用n≥2的字符n-gram特征,确保特征空间的简洁性与可解释性。实验结果表明,在严格控制泄露的闭集设置下,尽管未达到最优准确率,但该方法展示了在极简假设下字符级证据所能达到的性能边界,验证了其在透明性与安全性要求下的有效性。

链接: https://arxiv.org/abs/2607.22884
作者: Sanda-Maria Avram,George C. Ţurcaş
机构: Babeş-Bolyai University (巴贝什-博耶伊大学)
类目: Computation and Language (cs.CL)
备注: 17 pages, 12 tables

点击查看摘要

Abstract:We propose CHiPS, a lightweight character-level authorship attribution method for Romanian texts. All reported experiments are closed-set: the true author is one of the candidate authors in the training data. CHiPS studies two complementary fingerprints of writing style: CH-SVM, a character-histogram classifier based on one-character marginal distributions, and FFT12-LR, a positional-signal classifier that represents selected characters and punctuation classes as impulse trains (binary indicator sequences over character positions) and extracts Fourier/Welch spectral descriptors. We also report CHiPS-F, a leakage-safe decision-level fusion variant, and an optional top-5 listwise reranker trained only on out-of-fold predictions. The method requires no tokenization, syntactic analysis, pretrained language model, or transformer fine-tuning, and it avoids character n -gram features with n \geq 2 in the histogram component. On a locked grouped ROST split comprising 400 files from 392 source-text groups, written by 10 authors, with source-text-level evaluation and grouped five-fold model selection, CHiPS-F reaches 0.9310 accuracy and 0.9341 macro-F1. A matched but unrestricted character 2–5-gram TF–IDF SVM comparator reaches 1.0000 accuracy and macro-F1 on the same held-out groups, so the contribution is not a claim of best possible classification accuracy. Instead, the experiments ask how far restricted, transparent character evidence can go under strict leakage control. On ROSTories-cleaned, a secondary ROST-overlapping corpus comprising 1,248 files from 1,240 source-text groups, written by 19 authors, the same protocol gives 0.8919 accuracy and 0.8708 macro-F1 for CHiPS-R.

[NLP-100] PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs

【速读】: 该论文旨在解决低资源语言(如孟加拉语)在数学应用题(Mathematical Word Problems, MWPs)研究中因缺乏大规模标注数据集而面临的瓶颈问题。现有研究多集中于高资源语言,导致孟加拉语等语言的自然语言理解与量化推理能力评估严重不足。为此,本文提出PatiGonit22K——一个扩展后的孟加拉语数学应用题数据集,包含22,441道题目,通过在原始PatiGonit数据集基础上引入大量复杂数学问题进行扩充。其解决方案的关键在于:构建一个规模更大、难度分级均衡、涵盖单步与多步运算的高质量数据集,并确保每道题目经过精准翻译、标注、文化适配及数学正确性验证,从而为低资源语言中的数学推理与教育型自然语言处理(NLP)研究提供可靠且全面的基准支持。

链接: https://arxiv.org/abs/2607.22859
作者: Swastika Kundu,Azizul Hakim Fayaz,Tashreef Muhammad
机构: Ahsanullah University of Science and Technology, Dhaka, Bangladesh; Southeast University, Dhaka, Bangladesh
类目: Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning. Despite recent progress in high resource languages, Bengali remains underexplored due to the limited availability of large scale annotated datasets. In this work, we introduce PatiGonit22K, an expanded Bengali MWP dataset containing 22,441 problems, developed by extending the original PatiGonit dataset with a substantially larger collection of complex mathematical problems. The dataset includes both simple and multi operation equations, providing a balanced benchmark for evaluating mathematical reasoning across different difficulty levels. Each problem is carefully translated, annotated, culturally adapted, and verified to ensure linguistic consistency and mathematical correctness. By increasing both the scale and complexity of Bengali MWPs, PatiGonit22K provides a more comprehensive resource for future research on mathematical reasoning and educational NLP applications in low resource languages.

[NLP-101] Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

【速读】: 该论文旨在解决组织在将语言模型适配至内部应用场景时面临的两大挑战:一是提升模型在特定领域任务上的性能,二是应对敏感数据隐私泄露的风险。传统方法如开源模型微调(fine-tuning)或手动提示优化(prompt optimization)往往操作复杂且资源消耗大。为此,本文提出一种基于API层控制的轻量级替代方案——通过用户定义的向量对模型解码过程中的logits进行偏置(logit bias),实现无需修改模型权重、也不依赖梯度信息的模型适应。其核心创新在于设计了一种黑箱学习方法,用于学习一个全局固定的、上下文无关的logit偏置向量,并在每次解码步骤中添加该向量。该方法基于KL正则化的强化学习目标,推导出一种闭合形式的反倾向估计器(inverse-propensity estimator),利用轨迹(rollouts)、奖励信号和词元概率实现无梯度优化。实验表明,该方法在数学推理与逻辑推理基准测试中显著优于基础模型,同时所需的可训练参数远少于传统微调方式。结果表明,学习到的logit偏置是一种在极低访问权限要求下实现高效模型适配的轻量化机制。

链接: https://arxiv.org/abs/2607.22837
作者: Ofek I. Cohen,Lior Shani,Aviv Rosenberg,Ankur Samanta,Tal Wagner,Yonathan Efroni
机构: Google Research; Tel Aviv University
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 36 pages, 4 figures, 6 tables. Appendix included. Code available at this https URL

点击查看摘要

Abstract:Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. However, such adaptation remains non-trivial: it often requires operationally challenging fine-tuning of open-source models or ad hoc prompt optimization. We study a minimal alternative based on a simple API-level control: allowing users to bias the model’s logits with a user-defined vector. We develop a black-box method for learning a single context-independent logit-bias vector, added at every decoding step, without modifying model weights or requiring gradients. Starting from a KL-regularized reinforcement learning (RL) objective, we characterize when such a fixed logit-bias vector can approximate the optimal prefix-dependent correction and derive a closed-form inverse-propensity estimator from rollouts, rewards, and token probabilities. Empirically, this simple decoding-time intervention improves over base models on mathematical and reasoning benchmarks while using far fewer trainable parameters than conventional fine-tuning. Our results suggest that learned logit bias is a lightweight mechanism for adapting language models under minimal access requirements.

[NLP-102] he Best Programming Language for Tokenmaxxing: An Investigation of Coding Agent Behavior Across Programming Languages

【速读】: 该论文旨在解决多语言编程代理(coding agents)在不同编程语言中存在显著的令牌消耗差异问题,揭示了当前主流模型在跨语言任务中效率不均的现象。研究发现,即使在控制问题难度的前提下,同一类任务在如Python、Java、Rust和OCaml等语言上的令牌使用量仍存在巨大差异,且这种差异在多个模型间具有高度一致性。其解决方案的关键在于通过精细化分析代理的执行轨迹(trajectory),将每个中间解的执行结果抽象为测试通过/失败的向量序列,并标注相邻解之间的修正行为;同时对轨迹中的文本内容进行语义分析,发现代理在陌生语言中频繁生成无法编译的代码,反复修改已通过测试的方案,倾向于在代码注释中规划逻辑,依赖自创输入而非官方测试用例,并常通过在Python中原型设计来规避对目标语言的直接处理。这些发现表明,基于语言的令牌效率(by-language token efficiency)应成为评估与开发多语言编程代理的重要指标,也为优化资源分配提供了依据。

链接: https://arxiv.org/abs/2607.22807
作者: Zixuan Wu,Carolyn Jane Anderson,Arjun Guha
机构: Northeastern University(东北大学); Wellesley College(威尔斯利学院)
类目: oftware Engineering (cs.SE); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Although coding agents are now very effective in a variety of programming languages, this paper first shows that the cost (in tokens) can very significantly by programming language. We evaluate five recent models on programming problems in Python, Java, Rust, and OCaml. We carefully control for problem difficulty, and show that there can be stark variation in token consumption that is consistent across models. To understand why, we analyze both the structure and content of agent trajectories. First, we re-execute every intermediate solution and abstract each trajectory as a sequence of test-outcome vectors, then label the work between successive solutions. This reveals agents repeatedly producing noncompiling solutions in unfamiliar languages and revising solutions that already pass. Second, we analyze trajectory text, finding that agents plan solutions in code comments, distrust the provided tests in favor of inputs they invent, and sidestep unfamiliar target languages by prototyping in Python. Our results show that by-language token efficiency is a metric that should be considered when benchmarking and developing multilingual agents, and, for the tokenmaxxer, a guide to the most expensive language to work in. Subjects: Software Engineering (cs.SE); Computation and Language (cs.CL) Cite as: arXiv:2607.22807 [cs.SE] (or arXiv:2607.22807v1 [cs.SE] for this version) https://doi.org/10.48550/arXiv.2607.22807 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-103] Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training

【速读】: 该论文旨在解决深度学习在自动抑郁检测中因个体差异导致的领域偏移(domain shift)问题,从而限制模型泛化能力的挑战。其核心解决方案是提出首个不依赖特定患者的多模态抑郁检测框架,通过引入领域泛化(Domain Generalization, DG)机制,联合利用声学与文本模态信息。关键创新在于采用双向长短期记忆网络(BiLSTM)结合模态内与跨模态注意力机制,并引入段级融合进行决策;同时,借鉴领域对抗训练神经网络(DANN)思想,在模型中加入梯度反转层(gradient reversal layer),以抑制模型对说话人身份的判别能力,从而学习更具域不变性的特征表示,有效缓解患者特异性偏差。实验基于Androids-Corpus数据集,采用五折交叉验证,在30秒分段时长下,最优组合为梅尔频谱(MelSpec)与ItalianBERT,引入DG后准确率提升2.5%,F1分数提升3.3%,达到93.2%准确率、93.2%精确率、96.2%召回率和94.2% F1分数,显著优于现有基准。大量消融实验证明了多模态融合、深层网络结构设计及领域泛化机制的协同贡献,共同提升了模型的鲁棒性与泛化性能。

链接: https://arxiv.org/abs/2607.22794
作者: Ali Tabaraei,Federico Simonetta,Stavros Ntalampiras
机构: University of Milan (米兰大学); Gran Sasso Science Institute (格兰萨索科学研究所)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)
备注: 12 pages, 8 figures, 6 tables. Accepted for publication in IEEE Transactions on Neural Networks and Learning Systems (TNNLS)

点击查看摘要

Abstract:Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability. To address this critical issue, we present the first patient-independent multimodal depression detection framework that incorporates domain generalization (DG), jointly leveraging both acoustic and textual modalities. The proposed model integrates bidirectional Long Short-Term Memory (BiLSTM) with intra- and cross-modal attention mechanisms, accompanied by segment-level fusion for decision-making. Generalization is further enhanced by applying a gradient reversal layer inspired by Domain-Adversarial Training of Neural Networks (DANN), which promotes domain-invariant representations by adversarially limiting the model’s ability to identify individual speakers, effectively reducing patient-specific bias. Conducting experiments on the Androids-Corpus dataset with a 5-fold cross-validation (CV) protocol, various pairings of audio and text feature extractors were evaluated over different segment durations, determining MelSpec and ItalianBERT as the optimal baseline at a 30-second segment duration. The addition of DG to this baseline yields a 2.5% increase in accuracy and 3.3% in F1-score, achieving 93.2% accuracy, 93.2% precision, 96.2% recall, and 94.2% F1-score, surpassing all existing benchmarks. Extensive ablation studies assess the impact of multimodal fusion, deep architectural choices, and DG, highlighting their combined contribution to robust and generalizable depression detection.

[NLP-104] Open Your Models Eyes: Video and Context-Aware Multimodal Backchannel Prediction ACL2026

【速读】: 该论文旨在解决现有回声信号(backchannel)预测方法仅依赖音频与文本模态而忽略视觉信息及更广泛对话上下文的问题,导致对复杂回声(如共情表达)的识别准确率受限。其核心解决方案是提出一种上下文感知的多模态对齐框架——CAMA-BC,通过分阶段的多层多模态对齐(Multi-Layer Multimodal Alignment, MMA)机制实现。关键在于:第一阶段“上下文对齐”(MMA-CA)利用无标注视频对话数据捕捉动态对话语境;第二阶段“回声对齐”(MMA-BA)在此基础上进一步优化表示以专门服务于回声预测任务。该方法显著提升了对复杂回声行为的识别性能,尤其在共情类回声的检测中表现出色。

链接: https://arxiv.org/abs/2607.22729
作者: Min-Jae Kim,Jun-Yeong Moon,Mujeen Sung,Gyeong-Moon Park
机构: Korea University, Seoul, Republic of Korea; Kyung Hee University, Yongin, Republic of Korea
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注: 18 pages, Accepted at ACL 2026

点击查看摘要

Abstract:Backchannels, which signal listener states like empathy and understanding, are fundamental to natural human interaction. However, current approaches rely solely on audio and text. This omits crucial visual cues, such as facial expressions and gestures, as well as broader conversational contexts, which are necessary for accurate prediction. In this paper, we introduce Context-Aware Multimodal Alignment for Backchannel Prediction (CAMA-BC), a novel framework that leverages visual information through Multi-Layer Multimodal Alignment (MMA). Our alignment process comprises two stages. First, Context Alignment (MMA-CA) utilizes unlabeled dialogues with videos to capture conversational contexts. Next, Backchannel Alignment (MMA-BA) fine-tunes the representations specifically for backchannel prediction. Experimental results show that CAMA-BC significantly outperforms both existing methods and simple multimodal baselines, with particular effectiveness in recognizing complex backchannels such as empathy.

[NLP-105] CausalGate: Causal Importance Distillation for Transformer Module Pruning

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在自适应推理中依赖观测启发式方法(如隐藏状态相似性或激活幅值)进行冗余模块剪枝时,难以捕捉对语义准确性至关重要的非线性结构计算的问题。现有方法基于相关性度量,往往忽略深层网络中细微但关键的因果结构关系,导致性能下降。为此,论文提出CausalGate——一种基于干预引导的计算高效Transformer推理框架。其核心解决方案是:在校准阶段,通过系统性地隔离并置零每个Attention与MLP子层的输出,利用最终logit分布的KL散度精确量化各层对语义任务的贡献,从而构建出具有因果意义的结构重要性层级。为消除运行时路由开销,该重要性层级被压缩为一组全局静态、轻量级标量门控参数,采用指数移动平均平滑目标与可微分成对排序损失进行优化。实验表明,CausalGate在TinyLlama-1.1B、Qwen2.5-3B和Llama-3.1-8B等多个模型上,在语言建模与常识推理基准测试中均显著优于主流动态路由与层跳过基线方法,实现了理论计算节省向实际硬件延迟降低的有效转化,且无额外运行时开销。

链接: https://arxiv.org/abs/2607.22720
作者: Kiran Nair,Smriti Regmi,Rodrigue Rizk
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
备注: In-Review at a conference

点击查看摘要

Abstract:Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules. However, these correlation-based metrics often fail to capture subtle, non-linear structural computations vital for semantic accuracy. We introduce CausalGate, an intervention-guided framework for compute-efficient transformer inference. During a calibration phase, CausalGate isolates individual Attention and MLP sub-layers, zeros out their respective outputs, and measures the exact semantic damage via the Kullback-Leibler divergence of the final logit distribution. To eliminate runtime routing overhead, this structural importance hierarchy is distilled into a global set of static, lightweight scalar gates using an Exponential Moving Average smoothing objective paired with a differentiable pairwise ranking loss. Evaluated on TinyLlama-1.1B, Qwen2.5-3B, and Llama-3.1-8B across language modeling and commonsense reasoning benchmarks, CausalGate consistently outperforms prominent dynamic routing and layer-skipping baselines, translating theoretical compute savings into concrete hardware latency reductions with zero operational overhead.

[NLP-106] RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus

【速读】: 该论文旨在解决生成式图像中隐含性别歧视(misogyny)内容的自动化检测难题,尤其针对互联网模因(meme)因视觉与文本模态间语义冲突、仇恨意图隐含且高度依赖文化背景而带来的挑战。其核心解决方案是提出GeoMVC(Geometric Interaction and Multi-View Consensus)框架,关键在于引入几何交互层(Geometric Interaction Layer),通过冻结的视觉与文本嵌入之间的哈达玛积(Hadamard product)和余弦相似度建模跨模态对齐,克服传统静态特征拼接的局限性;同时,采用多视角一致性策略(Multi-View Consensus),融合原始文本、长度过滤后文本及英文翻译文本三种视图的预测结果,有效缓解由噪声光学字符识别(OCR)和混合语言转写(code-mixed transliteration)引起的分布偏移问题。该方法在ICMI 2026 CC-MMD挑战赛中取得优异表现,但在泰米尔语等德拉维达语系和中文语境下的局部转写与混合语风讽刺建模方面仍存在显著挑战。

链接: https://arxiv.org/abs/2607.22709
作者: Md. Ajwad Hossain
机构: Chittagong University of Engineering and Technology (CUET)(吉大港工程与技术大学); Bangladesh(孟加拉国)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted for the CC-MMD Grand Challenge at the 28th ACM International Conference on Multimodal Interaction (ICMI 2026)

点击查看摘要

Abstract:The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny. Memes often rely on a semantic clash between visual and textual modalities, where hateful intent is implicit and culturally grounded. This paper presents GeoMVC (Geometric Interaction and Multi-View Consensus), developed for the CC-MMD Grand Challenge at ICMI 2026. To address the limitations of static feature concatenation, a Geometric Interaction Layer is proposed that models cross-modal alignment via Hadamard products and cosine similarity between frozen visual and textual embeddings. We further mitigate distribution shifts caused by noisy OCR and code-mixed transliteration through a Multi-View Consensus strategy, aggregating predictions across raw, length-filtered, and English-translated text views. The system achieved Rank 2 in the Malayalam partition (Macro F1: 0.892) and Rank 3 in the Chinese partition (Macro F1: 0.895) on Task A, while securing Rank 5 in the Tamil partition (Macro F1: 0.521). A detailed error analysis on the development partition highlights open challenges in modeling localized transliteration and code-mixed sarcasm across Dravidian and Chinese cultural contexts.

[NLP-107] Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

【速读】: 该论文旨在解决后训练阶段自动化人工智能研究中模型参数与运行时支撑框架(harness)之间存在的优化脱节问题。现有方法通常在固定架构下训练模型,即提示词、工具、技能、中间件和记忆机制等均保持不变,而将数据生成过程排除在优化目标之外,导致模型更新与决定研究轨迹质量的静态支撑结构之间产生不匹配。其解决方案的关键在于提出Co-Harness框架,实现模型参数与运行时支撑框架的联合优化。该框架通过交替进行支架优化与模型优化:利用基于大语言模型(LLM)的HarnessCritic分析失败的研究轨迹,识别支架层面的失效模式,并提出经验证的局部改进方案;随后,模型在由优化后支架生成的高质量轨迹上进行微调,从而将有效的支撑结构“蒸馏”进模型参数中。200小时以上的自主案例研究进一步表明,Co-Harness能够自主恢复系统崩溃、提升推理效率并发现集成策略,无需人工干预。这证明了联合优化模型与支撑框架是超越传统固定支架后训练的有效路径。

链接: https://arxiv.org/abs/2607.22688
作者: Zhengyu Chen,Teng Xiao,Huaisheng Zhu,Yige Yuan,Luan Zhang,Jingang Wang
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from. Existing pipelines typically train models under a fixed harness, including prompts, tools, skills, middleware, and memory, while leaving the data-generating process outside the optimization objective. This creates a mismatch between model updates and the static scaffolding that determines trajectory quality. We introduce Co-Harness, a framework that jointly optimizes the agent harness and model parameters during post-training. Co-Harness alternates between harness optimization and model optimization. An LLM-based HarnessCritic analyzes failed trajectories, identifies harness-level failure modes, and proposes validated local updates. The model is then fine-tuned on high-quality trajectories generated by the improved harness, distilling effective scaffolding into model parameters. A 200+ hour autonomous case study further shows that Co-Harness can recover from system crashes, improve inference efficiency, and discover ensemble strategies without human intervention. These results suggest that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.

[NLP-108] Imprompt: A Language Framework for Prompt Programming

【速读】: 该论文旨在解决现有提示编程(Prompt Programming)框架中存在的复杂性与不优雅问题,这些问题使得实际应用中难以有效描述复杂任务。其核心挑战在于现有方法往往将任务描述与底层执行细节耦合,导致可读性差、可维护性低。本文提出Imprompt——一种新型提示编程语言框架,其解决方案的关键在于重构提示程序的设计范式:强调提示程序应仅包含任务描述,并与底层执行细节解耦,从而实现更高层次的抽象与可编程性。进一步地,作者将结构化提示(Structured Prompting)定义为提示编程与提示程序“编译”的结合,通过形式化定义两类编译器来验证该理念;同时引入类型系统,建立类型检查与约束解码(Constrained Decoding)之间的对应关系,以增强程序的正确性与可控性。最终,通过实现编译器与类型检查器并在多个案例研究中评估,验证了该框架在提升提示编程可表达性、可维护性与可靠性方面的有效性。

链接: https://arxiv.org/abs/2607.22683
作者: Chentian Wu,Shengyuan Yang,Adithya Murali
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Programming Languages (cs.PL)
备注: 28 pages, 24 figures, 2 tables

点击查看摘要

Abstract:With the unprecedented success of Language Models (LMs), the science of Prompt Engineering has evolved the powerful idea of Prompt Programming, where prompts are treated as a programmable control surface for describing complex tasks and leveraging LM capabilities. However, existing prompt programming frameworks suffer from various complexities and inelegances, which make them hard to utilize in practice for effectively describing tasks. We propose Imprompt, a new language framework for the study and practice of prompt programming. We undertake a foundational investigation of prompt programming, and contend that prompt programs must contain only the task descriptions and must be decoupled from lower-level ‘execution’ details. We further develop this position by illustrating structured prompting as a combination of prompt programming and prompt program ‘compilation’. We exemplify this view by formally defining two compilers for Imprompt programs. We then explore the idea of typing for prompt programs and draw a correspondence between type checking and constrained decoding. Finally, we implement our compilers and type checkers and evaluate them on a variety of case studies. We believe our work contributes programming-language foundations toward the emerging area of prompt programming.

[NLP-109] From peer review nuances to best practices

【速读】: 该论文旨在解决同行评审数据中存在三个关键细微差异——论文版本(paper version)、评分版本(score version)以及输入格式(input format)所带来的数据异质性问题。这些差异可能对下游任务(如模型训练、评估与分析)产生显著影响,但其具体影响机制尚未被系统研究。论文通过系统刻画各类变体之间的差异,并量化其对下游任务性能的影响,揭示了数据不一致性可能引入的偏差与噪声。解决方案的关键在于建立一套标准化的数据处理与使用规范:对于数据提供方,应明确标注数据版本与格式信息,确保可追溯性与透明度;对于数据使用者,则需在数据预处理阶段识别并统一不同版本间的差异,采用基于验证的清洗策略以提升模型鲁棒性。这一方法论为构建高质量、可复现的同行评审数据分析体系提供了实践指导。

链接: https://arxiv.org/abs/2607.22681
作者: Sheng Lu
机构: TU Darmstadt (达姆施塔特工业大学); Google(谷歌); Meta(元); Stability.AI(稳定人工智能); Anthropic(Anthropic); Character.ai(字符人工智能); Claude(克莱德); OpenAI(开放人工智能); Qwen(通义千问)
类目: Digital Libraries (cs.DL); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:This report studies three nuances in peer review data: paper version, score version, and input format. We characterize how the variants differ, and measure their impact on downstream tasks. Based on our findings, we offer best practices for both data providers and data users.

[NLP-110] How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

【速读】: 该论文旨在解决大语言模型在后训练(post-training)过程中对多维度对齐(alignment)影响不明确的问题,尤其关注任务适应是否会导致模型原有对齐属性(如安全性、事实性、社会危害规避等)发生不可预期的偏移。其核心解决方案在于系统评估多种典型任务适应方法——监督微调(SFT)、KL正则化SFT以及基于可验证奖励的强化学习(RLVR)——在15个涵盖安全、事实性、立场稳定性、社会危害、可控性和指令遵循性六大关键领域的对齐表现。研究发现,后训练并非均匀地重塑模型对齐状态:尽管RLVR在提升任务性能的同时仅引发微小但非零的特定指标偏移,而标准SFT则导致显著的跨领域对齐漂移;KL正则化通过增强参考模型锚定效应有效缓解了这一问题,但仍未能达到RLVR的对齐保持水平。进一步的表征层面分析表明,对齐相关表征的动态变化与行为漂移高度一致。因此,该研究揭示了任务适应本质上是一种对齐干预,强调应将多维度对齐评估作为后训练流程的标准组成部分。

链接: https://arxiv.org/abs/2607.22676
作者: James Elcock,William F. Shen,Xinchi Qiu,Nicholas D. Lane
机构: University of Cambridge (剑桥大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 21 pages, 7 figures (includes references and appendices)

点击查看摘要

Abstract:Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model’s pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.

[NLP-111] Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

【速读】: 该论文旨在解决生成式人工智能(Generative AI)在面对特定意识形态叙事时,可能无意识地复现或强化虚假信息框架的问题,尤其关注现有机器遗忘(machine-unlearning)算法在抑制此类行为上的有效性。其核心挑战在于,语言模型不仅会直接响应与目标叙事相关的提示,还可能在间接、对比或抽象等更隐蔽的层面再现这些叙事结构,从而导致“遗忘不彻底”或“反向恢复”的现象。为此,论文提出了一种基于上下文的评估协议——层级化叙事抑制评估(Level-based Evaluation of Narrative Suppression, LENS),通过直接、归因、对比和抽象四个层次系统测试模型对目标叙事的复现能力。解决方案的关键在于引入“抑制-崩溃效率”(Suppression-Collapse Efficiency, SCE)评分机制,该指标在奖励目标叙事抑制的同时惩罚输出质量下降,从而实现对模型遗忘效果的量化评估与最优检查点选择。实验结果表明,经优化后的模型检查点可有效降低叙事复现率,且抑制效果具备跨任务迁移性;同时发现一个关键副作用:在抽象提示下,模型可能重新激活与目标叙事相关的真实世界主体(如国家或组织),揭示了叙事结构在模型内部深层表征中的顽固性。这一发现表明,LENS不仅是有效的诊断工具,也为深入理解叙事遗忘的内在机制提供了方法论支持。

链接: https://arxiv.org/abs/2607.22657
作者: Viktoriia Makovska,George Fletcher
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing machine-unlearning algorithms can suppress this behavior. We introduce Level-based Evaluation of Narrative Suppression (LENS), a contextualization based evaluation protocol for testing target narrative reproduction across direct, attributed, contrastive, and abstract resistance levels. We evaluate two source-grounded narratives: one framing Russia’s war against Ukraine as forced by NATO expansion, and one framing the United States as exploiting or abandoning Taiwan. The experiments cover four near-12B multilingual instruction models: Lapa LLM, Gemma-12B, Qwen-14B, and TAIDE-Gemma. We introduce the Suppression-Collapse Efficiency (SCE) score as a checkpoint selection summary that rewards target-narrative suppression while penalizing degraded outputs. Our results shows that selected checkpoints can reduce narrative reproduction and suppression may transfer beyond direct forget prompts. We also report entity recovery as a separate side effect: abstract A/B/C prompts can cause models to recover the real-world actors associated with the target frame after unlearning. These findings demonstrate that LENS is a successful diagnostic protocol for both reporting and guiding the further study of the deeper structure of narrative unlearning. Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22657 [cs.CL] (or arXiv:2607.22657v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2607.22657 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[NLP-112] You Talkin to Me?: A Network Analysis of Gendered Speaker-Addressee Patterns in Film Screenplays

【速读】: 该论文旨在解决电影对白中性别化说话人-受话人关系的结构性偏见问题,超越传统仅关注“谁在说话”的研究视角,深入探讨“谁被说话”以及跨性别对话动态如何影响交流结构。其解决方案的关键在于构建并分析一个包含4,600个有向对白事件的手动标注数据集,结合网络分析、卡方检验、配对统计比较及参与度转移分析,在三个研究中揭示:男性角色在整体上占据主导地位,无论在发言者还是受话者角色中;跨性别对白虽在平均方向上呈现对称性,但在影片层面存在聚集现象;同性别对白倾向于分散对话注意力,而跨性别对白则引发更紧密的二元互动循环。研究结论指出,电影对白中的性别偏见并非仅源于发言时长或对话方向失衡,而是根植于对话本身的架构——通过排除特定性别个体作为受话人以及其在互动结构中的边缘化定位来实现。

链接: https://arxiv.org/abs/2607.22656
作者: Samin Khan,Camilla Griffiths,Shrikanth Narayanan,Dan Jurafsky,Sabyasachee Baruah
机构: Stanford University (斯坦福大学); University of Southern California (南加州大学)
类目: Computers and Society (cs.CY); Computation and Language (cs.CL)
备注: 22 pages, 5 tables, 3 figures

点击查看摘要

Abstract:Objective: This paper investigates the gendered structure of speaker addressee relationships in film dialogue, asking not merely who speaks, but who is spoken to and how conversational dynamics unfold across gender lines. Methods: Using a manually annotated dataset of 4,600 directed dialogue events from 38 film screenplays, we apply network analysis, chi squared tests, paired statistical comparisons, and participation shift analysis across three studies. Key Findings: Male characters dominate as both speakers and addressees corpus wide, even in scenes with more women; cross gender dialogue is directionally symmetric on average but clustered at the film level; and same gender turns diffuse conversational attention while cross gender turns produce tighter dyadic reciprocation. Conclusion: Gender bias in film dialogue operates through the architecture of conversation itself, through exclusion from interaction and structural positioning as addressees, rather than through speaking time or within conversation directional imbalance alone.

[NLP-113] STAIF: A Stage-wise Optimization for Complex Instruction Following

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在执行包含多个显式约束的复杂指令时,难以严格满足各项约束尤其是硬性约束(hard constraints)的问题。现有对齐方法(如直接偏好优化,DPO)通常依赖整体奖励信号进行优化,导致对个别约束的严格遵守被弱化,尤其在分布外或多重约束场景下表现不佳。为此,论文提出STAIF(Stage-wise Training for Alignment with Hard and Soft Constraints),其核心在于分阶段解耦主观性(软性,soft)约束与客观可验证(硬性,hard)约束的对齐过程:第一阶段采用基于多负样本的偏好优化,增强模型对软性约束的敏感度;第二阶段引入可验证奖励强化学习(Reinforcement Learning with Verifiable Rewards, RLVR),确保硬性约束得到严格遵守。为支持该方法,研究构建了STAINSTRUCT数据集,一个包含约3.1万条高质量中英文双语复杂多约束指令的数据集。实验结果表明,STAIF在多个代表性基准上均达到领先性能,并展现出良好的泛化能力。

链接: https://arxiv.org/abs/2607.22649
作者: Jian Hong,Chen Cheng,Quan Liu,Yuhao Chen,Enhong Chen
机构: University of Science and Technology of China (中国科学技术大学); iFLYTEK Research Group (科大讯飞研究院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 16 pages, 6 figures

点击查看摘要

Abstract:Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment methods, such as DPO, optimize holistic reward signals that often underemphasize strict satisfaction of individual constraints, particularly under out-of-distribution or multi-constraint settings. In this paper, we propose STAIF, a stage-wise optimization framework that decouples the alignment of subjective (soft) constraints from the optimization of objectively verifiable (hard) constraints. Stage 1 applies preference optimization with multiple negative samples to sharpen sensitivity to soft constraints, while Stage 2 applies Reinforcement Learning with Verifiable Rewards (RLVR) to enforce strict compliance with hard constraints. To support this method, we construct STAINSTRUCT, a high-quality bilingual (English, Chinese) dataset of approximately 31,000 complex multi-constraint instructions. Extensive analyses validate the design of STAIF and show state-of-the-art performance on representative benchmarks against strong baselines, as well as genuine generalization.

[NLP-114] Masked Distillation: Internalizing the Chain-of-Thought in Language Models

【速读】: 该论文旨在解决大型推理模型(Large Reasoning Models, LRMs)在推理过程中生成冗长的中间推理链(Chain-of-Thought, CoT)所导致的高延迟、高内存占用及高昂服务成本问题。尽管最终答案的正确性并不依赖于中间推理步骤的准确性,且推理链长度也并非问题复杂度的可靠指标,但现有模型仍需生成完整链条以完成推理。为应对这一问题,论文提出一种名为“掩码蒸馏”(masked distillation)的知识蒸馏框架,其核心在于将原本由中间推理令牌表达的计算过程内化到语言模型参数中,使学生模型能够直接生成答案(或仅需极短的中间步骤)。该框架的关键创新在于:学生模型仅基于问题预测解码结果,而教师模型则在给定问题及其自身完整推理链的条件下,对学生的输出提供反馈,从而引导学生学习从问题到答案的直接映射。该方法在两种设置下实现:一是自蒸馏(self-distillation),即同一模型在“思考模式”下作为教师,在“非思考模式”下作为学生;二是双模型蒸馏(dual-model),即由更大的推理教师监督一个独立的小型非思考学生模型。此外,通过控制学生被监督的中间令牌长度,该框架可在完全内化(学生仅输出答案)与无内化(学生输出完整推理链)之间进行插值,实现对知识内化程度的系统性探索。在GSM8K(小学算术)和Countdown(数字谜题搜索任务)两个推理领域上的受控实验验证了该方法的有效性。

链接: https://arxiv.org/abs/2607.22629
作者: Durgesh Kalwar,Vardhan Palod,Subbarao Kambhampati
机构: Arizona State University (亚利桑那州立大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though the final answer correctness is not causally related to the trace correctness and the trace length is not a reliable indicator of the problem complexity. This raises a natural question: can the computation expressed in these intermediate tokens be internalized into the parameters of a language model, enabling it to produce answers directly (or with much shorter intermediate traces)? We introduce \textitmasked distillation, a knowledge-distillation framework in which a student LLM is trained to predict only the solution tokens conditioned on the question, while a reasoning teacher provides feedback on the student’s responses after conditioning on the question and its own CoT trace. We instantiate this framework in two settings: (i) a \textitself-distillation setting, in which the same model serves as the teacher in thinking mode and as the student in non-thinking mode, and (ii) a \textitdual-model setting, in which a larger reasoning teacher supervises a separate smaller non-thinking student over the solution tokens. By treating intermediate tokens as a scaffold which reasoning models use to fit over the solution tokens, We additionally vary the length of intermediate-token scaffolding the student is supervised on, interpolating between full internalization (the student emits only the solution) and no internalization (the student emits the full trace before the answer). We evaluate the framework through controlled experiments on two reasoning domains: GSM8K (grade-school arithmetic) and Countdown (a number-puzzle search task).

[NLP-115] Learning When to Reason for Text-to-SQL via SFT and DPO

【速读】: 该论文旨在解决当前文本到SQL(Text-to-SQL)方法在处理简单查询时因强制进行多步推理(Chain-of-Thought, CoT)而导致的计算资源浪费与推理延迟过高问题。尽管现有方法依赖以推理为核心的范式在复杂基准上表现优异,但大量真实场景中的查询为简单的查找或聚合操作,无需复杂的推理过程。为此,本文提出AutoThinkSQL框架,通过在监督微调(Supervised Fine-Tuning, SFT)和直接偏好优化(Direct Preference Optimization, DPO)中引入自动思考机制,使模型能够动态判断查询复杂度:对简单查询跳过深度推理,仅执行必要操作;对复杂查询则激活完整的CoT推理流程。该方案的关键在于实现了推理行为与查询难度的自适应对齐,从而在保持甚至提升性能的同时,显著降低平均输出令牌数(分别减少24.6%和18.3%)与平均延迟(分别降低17.1%和11.5%),在Spider与BIRD基准上均优于最优基线。

链接: https://arxiv.org/abs/2607.22622
作者: Soohyuk Jang,Jiheum Yeom,Nohil Park,Sang Hun Kim,Yoonyoung Choi,Kiwook Bae,Sungroh Yoon
机构: Seoul National University (首尔国立大学); Samsung Electronics (三星电子)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 8 pages, 5 figures. Model checkpoints are available at this https URL

点击查看摘要

Abstract:Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inference-time overhead. However, a large fraction of real-world queries are simple lookups or aggregations that can be resolved without multi-step deduction, making forced reasoning wasteful. Thus, we propose AutoThinkSQL, a framework that integrates an auto-thinking mechanism into both Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) on Text-to-SQL. Our approach enables the model to dynamically bypass reasoning for simple queries while invoking deep CoT for complex queries. On Qwen3-Coder-30B-A3B, our method achieves consistent gains compared to the best counterpart baseline on both Spider and BIRD benchmarks while simultaneously reducing average output tokens by 24.6% and 18.3%, and average latency by 17.1% and 11.5% compared to CoT-only generation. Further analysis indicates that the model learns to align its reasoning decisions with query difficulty.

[NLP-116] Revitalizing Public Urban Places through Cultural and Political Memory: A Technological Approach with LLM s and Augmented Reality

【速读】: 该论文旨在解决数字时代背景下记忆、场所与身份之间复杂关系的阐释难题,尤其关注如何通过新兴技术重构人们对历史文化遗产的空间感知与情感联结。其核心问题是:在快速城市化与数字化进程中,如何有效保存并传递具有文化与历史意义的场所记忆,使其超越时间限制,持续影响当代社会的身份认同。解决方案的关键在于构建一个融合数字孪生(Digital Twin)与虚拟现实(Virtual Reality, VR)的技术框架,通过视觉叙事与语义图像搜索等手段,实现对历史景观与城市空间的动态重建与沉浸式体验。该方法不仅能够识别并强化具有记忆唤起功能的视觉元素,还支持跨时间维度的文化政治实体表达,从而推动公共空间的文化再生与可持续发展。研究强调以技术为媒介,实现对当下状态的韧性评估、未来演化的模拟以及过往历史的深度激活,最终形成一种连接过去、现在与未来的综合性文化记忆传播机制。

链接: https://arxiv.org/abs/2607.22613
作者: Lara Vartziotis,Tina Vartziotis,Valentin Keckeisen,Frank Beutenmueller,Martin Obstbaum,Sotirios Kotsopoulos,Kostas Moraitis
机构: 未知
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Presented here: this https URL

点击查看摘要

Abstract:This paper explores the intersection of memory, place, and identity, examining how new technologies, particularly Apple Vision Pro, can illuminate this nexus. Leveraging digital twins and virtual reality, it investigates how memory is woven into landscapes and urban environments of cultural and historical significance, identifying visual elements that evoke memory and heritage. Applications such as Apple Vision Pro can facilitate image extension to define place identity, informing viewers about cultural and political entities across timelines. Visual storytelling can showcase the evolution of landscapes and the preservation of cultural heritage, while Virtual Reality (VR) enables the recreation of historical landscapes and urban-scapes. This immersive approach invites users to transcend temporal boundaries and experience the past dynamically. Semantic Image Search can support research by uncovering images related to monuments, tradition, or cultural identity. This research introduces a methodology to connect digital twins and virtual environments with urban and non-urban landscapes to illustrate cultural, historical, and environmental sustainability. Central to this approach is defining the resilience of the current state, its future evolution, and the significance of the past. These technologies facilitate a historical and cultural embrace while evoking the feeling of returning to a specific place years later. The methodology outlines the integration of technologies needed to revitalize public urban places through cultural and political memory. Through these applications, this paper contributes to research on digital twins of spaces, urban transformation, and cultural heritage preservation. By offering insights into the relationship between memory, place, and identity in the digital age, it supports a deeper understanding of our collective past and its impact on the present. Comments: Presented here: this https URL Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) Cite as: arXiv:2607.22613 [cs.CY] (or arXiv:2607.22613v1 [cs.CY] for this version) https://doi.org/10.48550/arXiv.2607.22613 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Tina Nepheli Vartziotis [view email] [v1] Mon, 15 Jun 2026 08:48:03 UTC (933 KB)

[NLP-117] okengeist: Multi-Turn Attribution Tracing in Agent ic Conversations

【速读】: 该论文旨在解决大语言模型在多轮对话中生成响应时,其输出所依赖的上下文来源难以追溯的问题,尤其关注跨轮次依赖关系的非线性、分层结构如何在生成过程中传播。传统上下文归因方法采用单次遍历全量上下文的方式,仅能识别表层依赖,无法捕捉多跳(multi-hop)推理路径和深层依赖关系,导致“溯源崩溃”(provenance collapse)现象。本文提出多轮上下文归因(Multi-turn Context Attribution, MTCA)任务,目标是不仅识别与目标响应直接相关的前序对话轮次,还能揭示这些轮次自身对更早上下文的依赖关系。解决方案的关键在于提出Tokengeist框架——一种方法无关且可扩展的归因机制,通过将对话轮次建模为有向无环图(DAG),实现对依赖路径的递归遍历,从而完整恢复跨轮次的依赖链条。实验结果表明,现有扁平化归因方法在多跳依赖召回率上不足20%,而Tokengeist达到90%以上;同时,研究构建了MTCABench基准数据集,涵盖3,845个目标片段、665个多轮对话及深度达14的黄金标注溯源图,验证了递归式归因在揭示复杂对话依赖结构中的有效性。

链接: https://arxiv.org/abs/2607.22610
作者: Jessica Tang,Shraddha Barke,Sharad Agarwal
机构: University of Toronto (多伦多大学); Microsoft Research (微软研究院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies propagate across prior turns? Existing context attribution methods process the full context in a single pass, recovering surface-level dependencies but missing the layered, non-linear structure of real-world dialogues and multi-step reasoning tasks. We introduce multi-turn context attribution (MTCA): given a target span in a model response, the task of tracing attribution backward across turns to identify not only which prior turns were directly relevant, but also how those turns themselves depended on earlier context. We propose Tokengeist, an attribution-method-agnostic and scalable framework that recovers full dependency paths by casting attribution as a recursive traversal of a directed acyclic graph (DAG) over conversation turns. We will release MTCABench, a benchmark of 3,845 target spans across 665 multi-turn conversations, annotated with gold provenance graphs reaching depths of up to 14, across four dependency types. Across four open-weight models, flat attribution methods fail to recover multi-hop dependencies, achieving under 20% source recall, while Tokengeist reaches 90%. Our results reveal systematic failure modes of single-pass attribution – which we term provenance collapse – and motivate attribution methods that reason recursively across turns.

[NLP-118] MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models

【速读】: 该论文旨在解决多模态大语言模型(Multimodal Large Language Models, MLLMs)在推理过程中因关键值(Key-Value, KV)缓存内存开销随上下文长度线性增长而带来的性能瓶颈问题,尤其针对视觉令牌数量庞大导致的缓存压力。现有预填充阶段(prefill-stage)的KV选择方法依赖于预填充阶段的统计信息来估计KV重要性,隐含假设预填充阶段的查询表示能代表解码阶段的真实查询分布,但在多模态场景下,解码阶段的查询表现出远高于预填充阶段的方差,导致在严格缓存预算下KV重要性估计不稳定。这种不稳定性使得微小的排序误差可能误删语义关键的视觉令牌,进而严重损害模型的指称对齐(grounding)与推理能力。为此,论文提出MM-ShiftKV,一种无需训练、面向解码阶段感知且仅在预填充阶段执行的KV选择方法。其核心创新在于:通过构建方差扩展的查询代理(variance-expanded query proxies),在预填充阶段近似模拟解码阶段的查询行为,并基于这些代理所累积的注意力质量(aggregated attention mass)来评估提示中各KV的重要性,从而实现更鲁棒的缓存选择。实验表明,在多个多模态基准测试中,MM-ShiftKV在严格缓存预算下持续优于现有方法。

链接: https://arxiv.org/abs/2607.22586
作者: Jinsong Shu,Chenyang Wu,Zhongle Xie,Baokun Wang,Lidan Shou
机构: Zhejiang University(浙江大学); Ant Group(蚂蚁集团); The State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学); Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新区(滨江)区块链与数据安全研究院)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 19 pages, 11 figures

点击查看摘要

Abstract:Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a major bottleneck due to the large number of visual tokens. Recent prefill-stage KV selection methods estimate KV importance from prefilling statistics, implicitly assuming that prefilling-time queries are representative of those encountered during decoding. We show that this assumption breaks down in multimodal inference, where decoding-time queries exhibit substantially larger variance than prefilling-stage representations, leading to unstable KV importance estimation under tight cache budgets. As a result, small ranking errors can disproportionately discard semantically critical visual tokens and degrade grounding and reasoning performance. We propose MM-ShiftKV, a training-free, decode-aware and strictly prefill-only KV selection method. MM-ShiftKV approximates decoding-time query behavior during prefilling by constructing variance-expanded query proxies and estimates prompt KV importance based on their aggregated attention mass. Experiments on multimodal benchmarks demonstrate that MM-ShiftKV consistently outperforms existing methods under strict KV-cache budgets. Our code is available at this https URL.

[NLP-119] Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach

【速读】: 该论文旨在解决标准检索增强生成(Retrieval-Augmented Generation, RAG)流水线中仅依赖语义相似度进行文档排序,而忽视源出处可信度的问题。其核心挑战在于,低可信度来源可能因语义相关性较高而被优先召回,从而引入噪声或误导性信息,影响生成结果的质量与可靠性。本文提出的解决方案关键在于引入领域感知的源可靠性先验(source reliability priors),通过为不同来源类型(source type)分配预设权重因子λ(s),对原始检索得分进行重加权:score(q, d) = sim(q, d) × λ(s)。该方法在健康领域120篇文档的受控数据集上验证,相较于仅基于语义相似度的基线模型,使Precision@5从0.48提升至0.72,并在评估的对抗威胁模型下显著降低低可信度文档的检索频率。实验依托密尔沃基工程学院的高性能计算集群Rosie完成,确保了计算资源支持下的可重复性与可靠性。研究结果表明,引入源可信度先验可有效缓解RAG系统中因源质量下降导致的性能退化问题,具备一定的实际应用潜力。

链接: https://arxiv.org/abs/2607.22584
作者: Yuktha Tata Koganti,Hugo Garrido-Lestache Belinchon
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:

点击查看摘要

Abstract:Standard Retrieval-Augmented Generation pipelines rank retrieved documents by semantic similarity alone, without accounting for source provenance or credibility. This work evaluates a simple and interpretable modification to RAG retrieval ranking that incorporates domain-informed source reliability priors. Each document is assigned a prior lambda(s) based on its source type, and retrieval scores are reweighted using score(q, d) = sim(q, d) * lambda(s). The framework is evaluated against a similarity-only baseline on a 120-document health-domain corpus. In this controlled setting, source-aware reranking improves Precision@5 from 0.48 to 0.72 and reduces average adversarial document retrieval under the evaluated threat model, where low-credibility sources are identifiable via metadata. All experiments were executed on Rosie, the high-performance computing cluster at the Milwaukee School of Engineering, which provided the GPU-accelerated infrastructure necessary to run the full experimental pipeline reliably and reproducibly. These results suggest a potential mitigation strategy for source quality degradation in RAG pipelines, within the limits of the experimental setup described.

[NLP-120] Same Question Different Answers: Evaluating LLM Reliability Beyond Accuracy

【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在面对语义等价但表述不同的问题时,其推理与回答的可靠性问题。尽管模型在标准基准测试中表现出较高的整体准确率,但其对同一问题不同措辞版本的响应却存在显著不一致性,暴露出模型知识应用的脆弱性。研究发现,在事实问答与数学推理任务中,13个模型在4个基准上的实例级行为极不稳定:大量问题在不同改写形式下出现答案翻转现象,错误匹配率超过23%;尤其当仅以原始提问正确为条件时,答案翻转率更高,表明单一提示下的正确性并不能反映模型的真实可靠性。然而,研究同时观察到多数问题至少存在一种改写形式可使模型给出正确答案,说明模型具备潜在的完整知识,但知识检索过程受输入表述影响显著。为此,论文提出一种简单的自改写(self-paraphrasing)策略,通过生成多个语义等价的输入并聚合结果,有效恢复被遮蔽的潜在知识,从而在推理阶段提升性能。因此,该研究的关键在于揭示标准准确率指标可能掩盖模型在语义不变性下的严重不稳定性,并主张引入对等价输入的一致性评估作为衡量模型可靠性的更优范式。

链接: https://arxiv.org/abs/2607.22554
作者: Kazem Faghih,Yize Cheng,Shoumik Saha,Mobina Pournemat,Armin Gerami,Soheil Feizi
机构: University of Maryland, College Park (马里兰大学学院帕克分校)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change under meaning-preserving paraphrases across factual question answering and mathematical reasoning tasks. Across four benchmarks and 13 models, we find that model outputs frequently depend on the exact wording of the prompt. While overall accuracy typically changes only modestly across paraphrases, instance-level behavior is far less stable: for many questions, models alternate between correct and incorrect answers depending on phrasing, with mismatch rates reaching more than 23%. Conditioning on questions that are answered correctly in their original form reveals even larger failures measured by answer flip rates, showing that single-prompt correctness is often a poor indicator of reliability. At the same time, we find that models often produce a correct answer for at least one paraphrase of a question, suggesting that the underlying knowledge is present but inconsistently retrieved. Building on this observation, we show that a simple self-paraphrasing strategy can partially recover this latent knowledge and improve performance at inference time. Together, these findings suggest that standard accuracy metrics can mask substantial instability, and that evaluating consistency across equivalent inputs provides a clearer picture of LLM reliability.

[NLP-121] Evaluating the Impact of Reviewer Guideline Design on LLM -Based Automated Peer Review ACL2026

【速读】: 该论文旨在解决科学界日益增长的同行评审工作负荷问题,探索如何通过自动化手段提升同行评审效率与质量。其核心挑战在于:如何设计有效的评审指导规则以支持生成式 AI 在自动化同行评审中的应用。研究的关键发现是,经过会议实践不断优化的正式会议评审指南(official conference guidelines)在生成与人类评审判断高度一致的评价结果方面表现最优,表明长期积累的实践经验所提炼出的评估标准对自动化评审具有显著指导价值;相比之下,基于大语言模型(LLM)模仿高质量人工评审生成的“仿审稿人”指南则普遍效果较差。此外,研究还揭示强制采用严格结构化评分量表会系统性降低评审性能,凸显了在自动化评审中保留主观性与整体性判断的重要性。因此,解决方案的关键在于:优先采用经实践验证的、非机械化的、具有灵活性的评审准则,而非完全依赖形式化或机器生成的规则。

链接: https://arxiv.org/abs/2607.22553
作者: Haowen Li,Yoichi Ishibashi,Masafumi Oyamada
机构: Keio University(庆应义塾大学); NEC Corporation(日本电气公司)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 18 pages, 2 figures, ACL 2026 Findings

点击查看摘要

Abstract:Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review. Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.

[NLP-122] MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities LREC2026

【速读】: 该论文旨在解决科学文献中数学表达式自动形式化为可执行符号代码(即公式形式化,Formula Formalization)所面临的高质量、领域专用标注数据集严重匮乏的问题。其核心挑战在于如何构建一个既能支持精准人工标注又具备可扩展性的标注框架,以推动该任务的数据基础设施建设。解决方案的关键在于提出MioFFAn——一个开源、文档中心化且可定制的标注框架,基于MioGatto架构进行扩展,引入了针对公式形式化的特定功能,如目标公式的选取与符号代码辅助生成;通过支持用户自定义符号分类体系和符号算子兼容性配置,确保框架在不同专业科学领域中的适应性;同时,该框架设计了模块化自动化流程,利用大语言模型(Large Language Models, LLMs)实现部分自动化,并通过严格输出格式规范支持研究人员迭代优化自动化策略,结合标准自然语言处理(NLP)指标评估不同方法的有效性,从而验证了“人机协同”(human-in-the-loop)范式在提升标注效率与质量方面的可行性。

链接: https://arxiv.org/abs/2607.22552
作者: Nicolas Sibuet,Horacio Saggion,Riccardo Rossi
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG); Software Engineering (cs.SE)
备注: Presented in the 3rd International Workshop on Natural Scientific Language Processing (NSLP 2026), co-located at LREC2026

点击查看摘要

Abstract:The automatic translation of mathematical expressions in scientific literature into executable symbolic code (a process we refer to as Formula Formalization) is hindered by a severe scarcity of high-quality, ground-truth datasets specialized for technical scientific domains. In this paper, we present MioFFAn, an open-source, document-centric, and customizable framework designed to facilitate rapid annotation for this task. Building upon the MioGatto architecture, we extend existing features to overcome structural limitations and pivot its scope by introducing specific functionalities for Formula Formalization, such as selection of equations of interest and aided symbolic code specification. By allowing users to configure custom taxonomies and properties for identified symbols, and compatible symbolic operators, we ensure the framework is adaptable to diverse specialized scientific fields. Furthermore, MioFFAn is designed to incorporate partial automation via Large Language Models. By defining a modular set of automated sub-tasks with strict output formats, we enable researchers to iteratively refine automation capabilities and evaluate competing strategies using standard NLP metrics. We specify the current automation methodology and perform a preliminary evaluation that demonstrates to efficacy of this human-in-the-loop approach.

[NLP-123] Explaining GAND: A Resource on Gender-Ambiguous Natural Data Contrastive Attribution

【速读】: 该论文旨在解决机器翻译(Machine Translation, MT)系统在缺乏明确性别线索时产生性别偏见的问题。当前MT系统常依赖默认行为与刻板印象进行翻译,导致用户在追求自我表达的背景下遭受误导或伤害。为深入理解此类系统在性别模糊场景下的翻译机制,亟需自然且真实的基准数据集。为此,本文提出GAND(Gender-Ambiguous Natural Data),一个专为分析上下文线索对翻译中性别指代影响而设计的英文源句基准数据集。其核心解决方案在于利用GAND开展可解释性分析:将部分GAND语句翻译为两种具有语法性别区分的语言,并辅以人工构建的对比翻译版本;通过特征归因分析,识别出影响目标语言中模糊指代实体性别判断的关键源词及其上下文语境,从而揭示模型决策背后的性别敏感因素。

链接: https://arxiv.org/abs/2607.22546
作者: Janiça Hackenbuchner,Jasper Degraeuwe,Arda Tezcan,Joke Daems
机构: Ghent University (根特大学); Language and Translation Technology Team (LT3) (语言与翻译技术团队)
类目: Computation and Language (cs.CL)
备注: Accepted at EAMT2026: Technical Track

点击查看摘要

Abstract:Machine translation (MT) systems continue to produce gender-biased translations. In a time where self-expression is paramount, mistranslations based on default behaviour and stereotyping can lead to harm for users of these systems. To better understand how these systems translate gender in the absence of clear gender cues, we need benchmarking resources that reflect gender-ambiguous scenarios in a natural way. To this end, we present GAND, a gender-ambiguous natural data benchmarking resource for MT consisting of English source sentences, specifically designed to analyse the influence of contextual cues on gender in translation. We leverage GAND to conduct an interpretability analysis: we translate a subset of GAND into two grammatical gender languages and extend these with manually crafted contrastive translations. A following feature attribution analysis reveals source words in context that inform the gender translation of an ambiguous referent entity in the target translation.

[NLP-124] Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

【速读】: 该论文旨在解决在金融服务业(BFSI)及代理型应用(agentic settings)中部署大语言模型时,现有开源安全防护机制无法在单次推理过程中同时有效应对提示注入(prompt injection)、通用危害(general harm)与监管合规性(regulatory compliance)三大核心安全挑战的问题。其解决方案的关键在于提出Semalith v1.4——一个参数量为184M的DeBERTa-v3-base多任务分类器,通过在一个前向传播中实现三轴协同安全分类:包含九类子类型的提示注入检测、通用危害识别以及十一类金融服务业特定合规标签判定。该模型采用22分类头与4分类辅助超类别头联合加权损失训练,基于从49个公开来源挖掘的76,204条去重数据(使用SHA-1去重并确保评估集零污染,最大污染率仅0.22%),在22个基准测试上表现出色,尤其在提示注入检测方面显著优于Llama-Guard-3-8B(在7个提示注入任务全胜,整体18项中胜出11项,参数量仅为后者的1/44,且在208个良性代理提示上实现0.000的假阳性率,远优于后者0.063)。尽管在通用危害评测中仍落后于Llama-Guard-3,但研究明确指出这一互补性差异,并在第6节披露了六个已知弱点,为实际部署提供清晰指导:若用于对话内容审核,推荐使用v1.3;若需高覆盖率的BFSI标签或零假阳性率,则应选用v1.4。

链接: https://arxiv.org/abs/2607.22545
作者: Tejasvi C. Addagada
机构: 独立研究者(Independent Researcher)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
备注: 16 pages, 8 tables, no figures

点击查看摘要

Abstract:Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail addresses in a single inference pass. Semalith v1.4 is a 184M-parameter DeBERTa-v3-base classifier performing simultaneous three-axis safety classification including prompt injection, general harm, and financial-services regulatory compliance, in a single forward pass. Its 22-class head (BENIGN, nine prompt-injection sub-types, general-harm, eleven BFSI labels) is trained with a 4-class auxiliary super-category head under jointly weighted loss, on a 76,204-row corpus mined from 49 public sources with SHA-1 deduplication against every held-out evaluation set, with 21 of 22 benchmarks at zero contamination (max 0.22%). Against Llama-Guard-3-8B on 22 held-out benchmarks, Semalith v1.4 wins every prompt-injection evaluation (7/7) and 11 of 18 benchmarks overall at 44x fewer parameters, with FPR = 0.000 on 208 benign agentic prompts vs 0.063 for Llama-Guard-3-8B. On general-harm benchmarks (WildGuardMix, HEx-PHI, HarmBench), Llama-Guard-3 leads; this complementary split is documented in Section 4. Six measured weak spots are disclosed in Section 6. Deployment guidance: v1.3 is recommended for conversational moderation deployments (ToxicChat F1 0.624); v1.4 is recommended when BFSI label coverage or zero-FPR on benign agentic prompts is the priority. Comments: 16 pages, 8 tables, no figures Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR) Cite as: arXiv:2607.22545 [cs.LG] (or arXiv:2607.22545v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.22545 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Tejasvi Addagada [view email] [v1] Wed, 6 May 2026 10:14:30 UTC (22 KB)

[NLP-125] Creative Integration: A Decidable Criterion of Creativity

【速读】: 该论文旨在解决“整合”(integrative)概念在学术与技术实践中缺乏明确定义和可操作判断标准的问题,尤其区分真正的创造性整合(creative integration, CI)与表面化的重新描述。其核心挑战在于:如何在不依赖主观评价的前提下,判定某一整合是否真正实现了对世界描述的压缩——即通过融合两个原本冲突或不兼容的元素(A与B),在固定描述语言下显著缩短整体描述长度(描述长度比 $ C = L_{\text{pre}} / L_{\text{post}} < 1 $)。论文的关键解决方案是提出一个可判定、可验证且具备排斥伪整合能力的判据体系:通过四个二元合取门(conjunctive gates)实现判断的可计算性,并构建一套伪整合类型学以识别并排除外观类似但本质非整合的案例。该判据通过四项可证伪的测试进行验证——独立计算检验、对困难负样本的区分能力、跨样本预测性能以及描述语言鲁棒性——全部获得显著通过。因此,该研究的核心贡献并非“创造力即压缩”的哲学主张,而在于将这一理念转化为一个可引用、可复现的基础性判断工具,从而为更广泛的创造性智能研究提供形式化基准。

链接: https://arxiv.org/abs/2606.13977
作者: Yoshinori Nomura
机构: Mirage Mountain Technologies(幻影山科技)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 18 pages, 1 figure

点击查看摘要

Abstract:“Integrative” solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration – one that makes the world cheaper to describe – from a tidy re-description. Building on the lineage that treats creativity and intelligence as compression, we give such a criterion for creative integration (CI): the resolution of a real conflict between A and B is CI if and only if, under a fixed description language, the description length strictly shrinks (C = L_pre/L_post 1), with the reduction located in the conflict itself. We make the judgment decidable through four binary, conjunctive gates, and we fix its extension through a taxonomy of pseudo-integration that names and rejects the look-alikes. We back the criterion with a curated, multi-domain corpus and – crucially – validate it not by human inter-rater agreement but by four falsifiable tests it could fail: an independent computational check, discrimination against hard negatives, out-of-sample prediction, and description-language robustness; all pass with margin. The contribution is not “creativity is compression” but its decidability, discrimination, and corpus: on this account, what makes a move genuinely creative – rather than merely novel – is that it compresses a conflict, with novelty and value as downstream symptoms; whether all creativity is so constituted we state as an explicit conjecture. We claim only the sign of C-1; we judge, not generate. The result is a citable primitive for a broader program. Comments: 18 pages, 1 figure Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI) ACMclasses: I.2.7; I.2.0; F.4.1 Cite as: arXiv:2606.13977 [cs.CL] (or arXiv:2606.13977v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2606.13977 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Yoshinori Nomura [view email] [v1] Thu, 11 Jun 2026 23:49:25 UTC (26 KB) Full-text links: Access Paper: View a PDF of the paper titled Creative Integration: A Decidable Criterion of Creativity, by Yoshinori NomuraView PDFHTML (experimental)TeX Source view license Current browse context: cs.CL prev | next new | recent | 2026-06 Change to browse by: cs cs.AI References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

信息检索

[IR-0] Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

链接: https://arxiv.org/abs/2607.24651
作者: Zhuchenyang Liu,Yao Zhang,Yu Xiao
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document. Under this interface, vision-language models often fail to identify the right regions even when the answer is correct, a failure known as Attribution Hallucination. We present a study that investigates whether this failure is partially limited by what the model can express through coordinates. On a verified bilingual CiteVQA subset, we compare the coordinate interface with a language interface in which the model outputs only text, quoting its evidence verbatim, and a multimodal retriever returns the location of each quote as a page region proposed by a layout parser (tables and figures are quoted through their captions or notes); the comparison is repeated over six open vision-language models. Compared with the coordinate interface, evidence recall rises from at most 8 points to between 26 and 47 and the hallucination rate roughly halves, with little change in answer quality. Building on this comparison, we use the same quote-and-retrieve pipeline as a training scaffold: because region-level evidence labels are expensive to collect for long documents, we introduce a GRPO recipe whose reward is a judge’s reading of the gold answer and crops of the retrieved regions, training the model to quote better evidence without any region labels and raising an 8B backbone’s strict attributed accuracy from 22.4 to 33.8. These findings indicate a practical path to improve attribution"without a coordinate interface and without costly region-level supervision.

[IR-1] LaRec: Unleashing LLM -based Latent Reasoning for Generative Recommendation

链接: https://arxiv.org/abs/2607.24617
作者: Yu Xia,Zihan Lin,Wei Yang,Rui Zhong,Cheng Chen,Huan Ren,Yao Hu
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) have shown great promise in recommendation due to superior reasoning abilities. However, existing methods mainly rely on explicit Chain-of-Thought (CoT), resulting in verbose reasoning texts and inefficient response times. latent reasoning aims to balance efficiency by thinking within a continuous latent space, yet it faces two major challenges: (1) Lack of Fine-grained Supervision: Latent reasoning relies solely on feedback from the final labels, providing sparse supervisory signals that struggle to effectively guide the optimization of multiple hidden reasoning steps. (2) Single Reasoning Path: The deterministic nature of latent reasoning impedes the exploration of users’ diverse interests and preferences, thereby limiting the recommendation capabilities of LLMs. To address these issues, we propose \textbf LaRec , an efficient generative recommendation framework designed to unleash the potential of latent reasoning in LLMs. LaRec consists of two core stages: First, we design Latent Pre-training that empowers LLMs with latent reasoning capabilities by providing rich supervisory signals to the latent space reasoning via step-level alignment and process direction alignment. Second, we introduce Personalized RL-tuning. Specifically, we construct a personalized Gaussian Mixture Distribution for each user based on their historical interests. By randomly sampling distinct reasoning starting points from this distribution during training, we guide the LLMs to traverse diverse reasoning paths within the latent space, enabling efficient exploration of user’s multi-faceted interests. Experiments on multiple datasets show that LaRec significantly outperforms existing baselines with comparable efficiency.

[IR-2] One Graph Multiple Gains: Single High-Quality Item-Item Graph for Multimodal Recommendation ACM-MM2026

链接: https://arxiv.org/abs/2607.24607
作者: Jinfeng Xu,Zheyu Chen,Ziyue Peng,Shuo Yang,Jinze Li,Zewei Liu,Shujie Li,Yipeng Du,Edith C. H. Ngai
类目: Information Retrieval (cs.IR)
备注: Accepted by ACM MM 2026

点击查看摘要

Abstract:Multimodal recommendation leverages item multimodal features alongside collaborative signals to capture user preferences. While item-item graphs have become a key building block in advanced models, existing methods typically construct them with noisy similarity edges and limit their role to a single function of item-item representation propagation, leaving substantial potential untapped. In this paper, we propose IIMRec, a framework that constructs a single high-quality item-item graph during preprocessing and systematically reuses it across three stages of the recommendation pipeline: representation enhancement, interaction graph enhancement, and optimization enhancement. The graph is built by fusing semantic and co-occurrence signals, then refined via Neighborhood Consistency Edge Reweighting (NCER), which applies the triadic closure principle to amplify structurally reliable edges and suppress spurious ones. Once constructed, the graph is leveraged in three complementary ways: (1) Item-item propagation with a Residual II Gate (RIG) that adaptively controls per-item absorption of semantic neighborhood signals for representation enhancement; (2) A content-guided UI graph expansion that introduces virtual user-item edges through high-confidence semantic neighbors for interaction graph enhancement; (3) II-Neighbor BPR Augmentation (INA) that treats top neighbors of positive items as discounted soft positives for optimization enhancement. We provide theoretical analysis showing that NCER reduces the spectral noise-to-signal ratio, RIG converges to a non-degenerate gating regime, and INA yields a tighter generalization bound. Extensive experiments on four datasets demonstrate that IIMRec consistently outperforms state-of-the-art baselines while running faster and consuming less GPU memory, with particularly strong gains under cold-start and sparse-interaction conditions. Comments: Accepted by ACM MM 2026 Subjects: Information Retrieval (cs.IR) Cite as: arXiv:2607.24607 [cs.IR] (or arXiv:2607.24607v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2607.24607 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[IR-3] DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing

链接: https://arxiv.org/abs/2607.24567
作者: Tobias J. Bauer,Christian Riess,Daniel Loebenberger,Christian Bergler
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
备注: 12 pages, 3 figures, 9 tables

点击查看摘要

Abstract:Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have gained extensive consideration in recent years. Deep learning-based methods offer better semantic capturing capabilities than traditional approaches relying on manual feature engineering. Moreover, they enable a data-driven approach to semantic hashing across diverse data modalities, yielding high-quality cross-modal hash codes within a shared Hamming space. Previous work investigated the properties of this Hamming space and introduced a loss function based on predefined so-called semantic channels with fixed width and Hamming distances derived from label similarities. However, this formulation also introduced discontinuities into the loss landscape, complicating optimization. Based on these observations, we propose a newly designed loss function, Dynamic Semantic Channel Hashing (DSCH), using dynamically sized and positioned semantic channels in order to avoid loss landscape discontinuities. Furthermore, we endorse the use of tie-aware Mean Average Precision (mAP) as evaluation metric as it addresses the ambiguity in sample retrieval ordering, which emerges from the discreteness of hash code distances. Finally, multiple experimental settings conducted on two popular datasets and incorporating two different model architectures provide strong evidence that training using the DSCH objective outperforms training using other state-of-the-art loss functions. In a total of 35 out of 40 cross-modal and intra-modal retrieval tasks, models trained with DSCH achieve significantly higher tie-aware mAP scores across all four tested hash code lengths, showing compelling results across model architecture and used dataset. The mAP score uplifts are consistent and amount up to 1.75 percentage points compared to the respective second best.

[IR-4] DeCoRAG : Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding

链接: https://arxiv.org/abs/2607.24554
作者: Shuo Wang,Kai Zhang,Wenyuan Huang,Yizheng Yu,Xia Liao,Junming Su,Qing Wang,Fang Xi
类目: Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV)
备注: 11 pages, 4 figures, 8 tables

点击查看摘要

Abstract:Advancing multimodal retrieval-augmented generation (RAG) for complex document understanding presents a formidable dual dilemma of accuracy and efficiency, particularly in graph RAG. Processing structurally sparse yet visually dense layouts, such as extracting a tiny data marker from a financial chart, often incurs computationally prohibitive token overhead while still triggering catastrophic hallucination. However, multimodal Graph RAG pipelines rely on graph-construction stages that assume Vision-Language Models (VLMs) can resolve sparse semantics within high-density layouts. We challenge this assumption, revealing that forcing VLMs to localize visual evidence, interpret semantics, and extract relations triggers a “Visual Attention Sink,” a mechanism driving catastrophic semantic loss, while full-page processing incurs massive computational overhead. Controlled interventions verify that this failure is boundary-driven rather than content-specific and that semantic anchoring mitigates it. To fundamentally correct this flawed paradigm, we introduce DeCoRAG, a multimodal Graph RAG pipeline that shifts knowledge processing from coupled visual-semantic reasoning to “Cognitive Decoupling.” Rather than passively processing raw pixels, its graph-construction stage establishes a macroscopic Semantic Anchor to neutralize the attention sink. This anchor subsequently drives our Region-Aware Pruning and Cropping (RAP-Crop) mechanism, shifting the reasoning space from dense, noisy backgrounds to purified, intent-driven semantic clusters. The resulting graph supports hybrid retrieval and answer generation. Across complex document benchmarks, DeCoRAG improves the semantic pass rate by up to 12.5 percentage points over the strongest baseline and generalizes to DocVQA. RAP-Crop reduces offline graph-construction prompt tokens by 40.8% without sacrificing end-to-end accuracy.

[IR-5] From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages

链接: https://arxiv.org/abs/2607.24542
作者: Th{é}otime de la Selle(ISC, HiSoMA, CNRS)
类目: Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient languages. Downstream semantic analysestext reuse identification, alignment, and semantic search-rely on sentence embeddings, yet existing methods transfer poorly to ancient languages: generic multilingual encoders underperform, specialized language models yield anisotropic representation spaces, and labeled similarity data is unavailable. We study two fully unsupervised strategies - TSDAE and contrastive sentence embedding (CSE) - that adapt a specialized token-level language model into a corpus-specific sentence encoder using only raw sentences. On the philologically central case of biblical reuse in patristic literature (2,935 expert-verified parallels in Latin and Ancient Greek, from Augustine, Jerome, and Athanasius), we decompose reuse identification into two separately evaluated tasks-binary detection and correspondence retrieval-and benchmark the adapted encoders against multilingual, specialized, distilled, and supervised fine-tuned baselines, as well as on artificially noised data simulating HTR artifacts and scribal abbreviations. The adapted encoders outperform all baselines on both tasks, with complementary profiles: TSDAE leads detection given a large in-domain corpus, while CSE leads retrieval, reaches its optimum with as few as 4-8k raw in-domain sentences-a few tens of seconds of training on a laptop GPU-and transfers across works and authors, including to noisy post-ATR text when retrained directly on it. UMAP atlases relate the geometric effect of each strategy to the measured gains, and the full pipeline-segmentation, fine-tuning, cross-corpus semantic search-is made available to non-specialists through the online tool Paraphrasis.

[IR-6] Robust Interpretation of Historical Documents in Knowledge Graphs Through Query Inference and Execution ICDAR2026

链接: https://arxiv.org/abs/2607.24475
作者: Sebastià Nicolau,Adrià Molina,Oriol Ramos Terrades,Josep Lladós
类目: Information Retrieval (cs.IR); Digital Libraries (cs.DL)
备注: Accepted at ICDAR2026

点击查看摘要

Abstract:The emergence of Large Language Models (LLMs) has redefined how users interact with information in digital environments. However, their widespread and often indiscriminate integration has raised significant concerns regarding reliability and trustworthiness issues that are particularly critical when accessing digital libraries and historical archives. How can one leverage the generalization capacity of an LLM without losing the level of accountability required for an archival institution? In this paper, we present an agentic retrieval system designed to deliver more accurate and verifiable access to historical data while preserving much of the flexibility associated with unconstrained LLMs. As a contribution to historical document analysis, we compare traditional Retrieval-Augmented Generation (RAG) with an agentic GraphRAG architecture in their ability to deliver historical information under realistic conditions, including the presence of OCR and transcription errors. We introduce a semi-symbolic framework that integrates word-spotting techniques for post-OCR correction with a knowledge graph representation that enables the agent to access information through synthesized queries. The interleaved collaboration between word spotting and code generation allows the agent to construct strong retrieval queries that are robust to misinterpretation and hallucination, while still leveraging approximate search when noise and uncertainty, common in historical document analysis, would otherwise hinder precise retrieval. Comments: Accepted at ICDAR2026 Subjects: Information Retrieval (cs.IR); Digital Libraries (cs.DL) Cite as: arXiv:2607.24475 [cs.IR] (or arXiv:2607.24475v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2607.24475 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[IR-7] Evaluating RAG for French immigration law: a benchmark and baseline study

链接: https://arxiv.org/abs/2607.24449
作者: Annia Abtout,Julien Delaunay,Monika Ewa Rakoczy
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:International recruitment in France requires navigating a layered legal framework absent from existing legal AI benchmarks. We present a publicly available benchmark and first comparative evaluation for this domain, covering permit-type recommendation, required-document retrieval, and legal citation coverage. Comparing a parametric LLM baseline against dense retrieval augmentation at two model scales (Qwen3.5-9B and -27B) on 52 annotated synthetic profiles, we find that retrieval improves administrative guidance at both scales, most notably permit-type accuracy. Our results confirm that retrieval grounding is important for more reliable administrative guidance in this domain, and motivate further investigation of hybrid retrieval strategies.

[IR-8] Unifying Generative Recall and Multi-Objective Ranking in a Single Decoder-Only Sequence

链接: https://arxiv.org/abs/2607.24439
作者: Ruochen Yang,Shuang Wen,Pengbo Xu,Yusheng Huang,Jiangxia Cao,Shuang Yang,Zhaojie Liu,Jiawei Sheng,Tingwen Liu
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Modern industrial recommendation systems typically separate recall and ranking into two independent stages. Although this cascade supports corpus-level retrieval and fine-grained multi-objective scoring, it causes objective inconsistency, information loss at the candidate hand-off, and redundant user-side context computation. Meanwhile, the generative recall and ranking scaling share a common Transformer-based modeling philosophy, where architectural consistency creates a natural opportunity for unified integration. However, direct sharing remains challenging since the two tasks require different information visibility and optimization methods. Therefore, we propose \textbfUniR ^2 , a \textbfUnified decoder-only Transformer that unifies Generative \textbfRecall and Multi-Objective \textbfRanking within a single heterogeneous sequence comprising user context, SID trajectory, and item features. Within this sequence, the generated trajectory serves as a representation bridge between recall and ranking, where Dual-Query Prefix-Causal Attention provides task-specific visibility. The two tasks share the base attention weights but retain separate optimization boundaries, with ranking-side LoRA preserving ranking adaptability without disrupting the generative backbone. Extensive offline experiments on large-scale industrial data demonstrate the effectiveness and efficiency of UniR ^2 for both recall and ranking. Long-term online A/B tests on Kuaishou platform further show consistent positive gains, validating the practicality of unified model in large-scale recommendation systems.

[IR-9] CORE: A Unified Cascaded Ordinal Relevance Estimation Framework for E-commerce Search

链接: https://arxiv.org/abs/2607.24417
作者: Zhi Jin,Xi Wang,Yunfei Li,Guojun Liu,Qingsong Hua,Wei Lin
类目: Information Retrieval (cs.IR)
备注: 11 pages, 5 figures

点击查看摘要

Abstract:Ranking relevance is a fundamental task in e-commerce search, directly affecting ranking quality and consumer experience. Although inherently an ordinal classification problem, it is commonly formulated as conventional multi-class classification, which overlooks the natural order among relevance levels and assigns equal penalties to adjacent and distant misclassifications. This mismatch leads to suboptimal learning objectives for practical relevance evaluation. To address this issue, we propose a unified cascaded binary classification framework applicable to both large language model inference and online BERT-based inference, which reformulates relevance estimation as a sequential decision process and decomposes multi-class prediction into a series of ordered binary judgments from higher to lower relevance tiers. For large language models, we design a step-wise reasoning procedure with pruning strategies and tier-specific reward functions. For the online BERT model, we replace the conventional classification head with multiple level-wise binary classifiers and distill the capabilities of large language models into the online model. Extensive offline industrial benchmark evaluations and online A/B experiments demonstrate that the proposed framework substantially improves relevance performance, reducing the online bad-case rate by 15.94%. Further analyses suggest that tier-wise modeling is effective for relevance estimation.

[IR-10] Occluded Oculus: Operationalizing Stylistic Obscurement

链接: https://arxiv.org/abs/2607.24411
作者: Robert Dilworth
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注: 46 pages, 16 figures, 5 tables

点击查看摘要

Abstract:What did it take for Hermes, the devout messenger of the Olympian gods, to slay Argus Panoptes, the multi-eyed giant of Greek myth? As the perfect guardian, Panoptes’ legion of ever-watchful eyes proved difficult – but not impossible – to defeat. The centerpiece of Hermes’ strategy was obfuscation and sabotage. Posing as a shepherd, Hermes sealed each of Panoptes’ eyes – eyes that would otherwise have alerted the fearsome giant to Hermes’ plot – and vanquished him. The moral of the story: when a challenger must surmount a formidable foe – one far greater in stature and vastly more equipped – crafty maneuvers are not merely advisable but indispensable for victory. In this work, the “challenger” is a collective leveraging adversarial tactics to overcome the “multi-eyed giant” of stylometric systems and surveillance apparatuses. To successfully claw back the privacy siphoned by the multi-eyed giant, the challenger must carefully evaluate their plan of attack, \textitTraceTarnish , and determine what does and does not work to anonymize the authorship of text. To that end, we conduct an ablation study of \textitTraceTarnish to better understand which module – Translation, Obfuscation, Imitation, or Injection – best confounds a stylometric system. Our results indicate that the most effective approach was Injection, meaning that inserting zero-width Unicode characters, homoglyphs, and intentional misspellings neutralizes the indefatigable eyes long enough to claim the head of the all-seeing giant.

[IR-11] CogRec: Structure-Cognitive Fast-and-Slow Reasoning for Generative Recommendation

链接: https://arxiv.org/abs/2607.24402
作者: Xiang Liu,Jingsong Su,Shuqi Zhao,Pengbo Mo,Yiming Qiu,Huimu Wang,Mingming Li,Jiao Dai,Jizhong Han,Songlin Hu
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Semantic-ID-based generative recommendation represents each item as a hierarchical discrete token sequence and reformulates next-item prediction as constrained sequence generation. Existing methods, however, mainly use Semantic IDs as target sequences to be memorized, leaving the hierarchy, intra-layer relations, and item neighborhoods underused as an explicit reasoning space. Explicit reasoning-enhanced generative methods often produce a natural-language rationale before the item identifier, but this rationale is only weakly coupled with the discrete SID space in which the final prediction is made. We propose CogRec, a structure-cognitive fast-and-slow reasoning framework that grounds intermediate reasoning in the same SID topology used for target generation. CogRec augments the vertical SID hierarchy with intra-layer semantic graphs and item-level neighborhoods, and introduces SID Routing to represent recommendation reasoning through layer-wise Match, LateralJump, and Explore operations. Exact matching implements fast semantic localization, whereas lateral and exploratory operations instantiate slower structural navigation. A supervised multi-stage pipeline aligns the newly introduced SID tokens, establishes direct SID generation, and trains natural-language and SID-routing reasoning branches from a shared checkpoint under the same trie-constrained output space. Experiments on three public sequential-recommendation benchmarks show that SID Routing improves its corresponding direct-generation, indicate that structure-grounded reasoning is most useful when prefix matching is insufficient but learnable SID-space transitions remain available, whereas long or weakly supported routes introduce additional decoding cost and accumulated errors. Code is available at this https URL

[IR-12] OxygenREC-v2: Internalizing Discrimination into Generative Recommendation

链接: https://arxiv.org/abs/2607.24255
作者: Guo Tang,Hanye Wu,Changjiang Han,Qingyang Li,Ming Zhang,Xiangyu Qian,Yanchen Qiao,Huanjie Wang,Zhi Ma,Zhen Li,Yaqiang Zang,Pinghua Gong
类目: Information Retrieval (cs.IR)
备注: 16 pages, 8 figures

点击查看摘要

Abstract:Generative recommendation unifies retrieval and ranking within a single model by autoregressively decoding semantic identifier (SID) sequences. Yet reliably incorporating behavior signals from clicks, cart additions, and orders remains challenging. Existing approaches either jointly optimize generative and discriminative objectives, requiring delicate trade-offs, or use a separate ranker as a post-hoc reinforcement-learning reward, risking out-of-distribution scoring and reward misalignment. We propose OxygenREC-v2, a generative recommender that Internalizes Discrimination into Generative Recommendation (IDGR). Rather than adding a separate discriminative objective, OxygenREC-v2 uses logged behavior to condition generation and supervise training. During pre-training, a behavior instruction conditions generation on the target behavior. During post-training, future interaction behaviors are exploited as privileged knowledge in our entropy-aware trajectory optimization self-distillation framework, enabling reward-model-free policy optimization. Throughout both training stages, OxygenREC-v2 maintains a single unified backbone. We implement OxygenREC-v2 as a 3B-parameter, 1B-activated MoE and deploy it on this http URL’s large-scale e-commerce platform. Across multiple online A/B tests, OxygenREC-v2 improves user click-through conversion rate (UCTCVR) by 1.6–4.4% and GMV by 2.8–6.8% over OxygenREC-v1.

[IR-13] A Model-Driven Pipeline for Data Quality Specification and Operationalization: A No-Code Approach for Domain Experts

链接: https://arxiv.org/abs/2607.24245
作者: Arno Kesper,Lukas Sebastian Hofmann,Markus Matoni,Gabriele Taentzer
类目: Information Retrieval (cs.IR)
备注: 10 pages, 12 figures

点击查看摘要

Abstract:High-quality data is essential for reliable analysis, decision-making, and research across domains. This is especially relevant in areas such as cultural heritage, where data is collected and curated manually, making it prone to quality issues like inconsistencies. To improve data quality, the data must be analyzed regularly using systematic quality analyses. Quality analyses validate the conformance of data to domain-specific expectations. These expectations are best understood by domain experts, who can express them using natural language. However, they rarely possess the technical expertise to formalize these expectations into executable quality analyses. Consequently, this process requires domain experts and data engineers, making it time-consuming and technically demanding. The required technical expertise and the resulting dependencies pose a significant challenge. To address this challenge, we present a pipeline for formalizing and operationalizing data quality constraints. We support this pipeline using QPM, a metamodel for defining templates for reusable quality analyses. The web application Constrainify enables tailoring templates to specific conceptual requirements and translating them into executable quality analyses via a tool-chain based on model-driven engineering subpipelines. The result is a set of reusable, repeatable, and domain-specific quality analyses.

[IR-14] Strategy-Aware Parameter-Efficient Adaptation for LLM -based Auto-Bidding

链接: https://arxiv.org/abs/2607.24232
作者: Songyue Cai,Lianyu Wang,Shan Gu,Ziru Xu,Jian Xu,Xiaofeng Zhu,Bo Zheng
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Advertising bidding has evolved from manual strategies to auto-bidding systems better adapted for large-scale, dynamic auction environments. While recent advances in Large Language Models (LLMs) offer strong reasoning for auto-bidding, existing methods suffer from shallow trajectory-text interactions and require costly fine-tuning, hindering the efficient use of pretrained knowledge under diverse constraints. To address these challenges, we propose SAGE, a novel Strategy-aware Auto-bidding framework Guided by LLMs for Efficient bidding. SAGE introduces a parameter-efficient multi-modal alignment framework for constrained auto-bidding with LLMs. Specifically, SAGE comprises three key components: (i) the position augmentation module adopts temporal-semantic positional embeddings to effectively capture the intrinsic dynamics and semantic structures; (ii) the text alignment module leverages gated cross-attention to align the embedding spaces of trajectory and text modalities, enabling effective multi-modal fusion while alleviating the computational overhead caused by long trajectories; (iii) the constraint-gated LoRA module employs constraints as routing signals, activating only a small subset of experts to adapt the behavior of a frozen LLM efficiently. Extensive experiments on large-scale auto-bidding benchmark demonstrate that SAGE consistently achieves superior performance while tuning less than 10% of the trainable parameters required by full fine-tuning. Ablation studies further validate the critical contribution of each component to the framework’s overall performance.

[IR-15] Integrating Factual and Normative Industrial Knowledge via Constraint-Aware Graph Attention for Process Plan Recommendation

链接: https://arxiv.org/abs/2607.24213
作者: Yuntong Chen,Yingqi Li,Yingying Xiao,Ziang Wang,Zewei Liu,Jiahao Liu,Xitian Tian,Lijiang Huang
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Integrating heterogeneous industrial knowledge, including factual relations and decision constraints, remains a core challenge in industrial information systems. Machining process planning exemplifies this problem because engineers must select operations by combining material properties, feature characteristics, and quality requirements. Existing methods rely mainly on similarity retrieval or classification, without a unified ranking objective or standardized evaluation. We propose PCA-GAT, which formulates machining process plan recommendation as a knowledge graph enhanced collaborative filtering problem. Bayesian Personalized Ranking provides the learning objective, while Recall@K and NDCG@K define evaluation. The knowledge graph supplies semantic structure when collaborative signals are sparse. Four domain constraints, material compatibility, precision requirements, feature applicability, and operation sequencing, are introduced as attention biases during graph propagation. Type-specific weights learn their importance, and an adaptive gate adjusts their influence using local context. On a real aerospace dataset with 115 parts and 507 plans, PCA-GAT achieves Recall@1 = 0.9087 and strong cold-start robustness, with about half the degradation of the strongest baseline under severe sparsity. Ablation studies show that knowledge graph enrichment is essential, constraints add value, and ungated constraint injection can hurt performance. The learned weights identify material-operation compatibility as the dominant factor, consistent with domain expertise. Results on three public benchmarks show no degradation when constraints are absent, supporting generalization beyond manufacturing. This study establishes a standardized recommendation protocol for engineering process planning and benchmarks seven methods across three categories, showing that knowledge representation is the main bottleneck.

[IR-16] Secrecy Energy Efficiency for IRS-Assisted Low-Altitude Communications: A D3QN-PER Based Approach

链接: https://arxiv.org/abs/2607.24183
作者: Ya Gao,Peina Zhao,Yiheng Li,Wenchi Cheng
类目: Information Theory (cs.IT); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:To address the security and energy efficiency challenges in low-altitude economy (LAE) wireless communications, we develop a secure synergistic network integrating unmanned aerial vehicle (UAV) and intelligent reflecting surface (IRS), with an emphasis on maximizing secrecy energy efficiency (SEE) for downlink transmission scenarios. In particular, firstly, we establish the channel transmission models for UAV-IRS assisted LAE communications network. Then, we formulate a non-convex fractional optimization problem for SEE maximization, involving three tightly coupled variables, i.e., the beamforming, IRS phase and UAV trajectory. To tackle the fractional structure and variable coupling, Dinkelbach’s method and equivalent transformations are leveraged to reformulate the objective function, which is then decoupled and decomposed into three independent subproblems via an alternating optimization strategy for iterative resolution. Slack variables and Semidefinite Relaxation (SDR) are further employed to convexify the subproblems of beamforming and IRS phase shift optimization, thereby obtaining their optimal solutions. For the UAV trajectory optimization subproblem, we propose a D3QN-PER algorithm, which integrates a Dueling Double Deep Q-Network with Prioritized Experience Replay, to tackle the slow convergence and training instability inherent in conventional Deep Q-Network (DQN). Numerical simulations validate the performance for our proposed joint optimization scheme. Comparative results demonstrate that the developed D3QN-PER-based algorithm outperforms existing state-of-the-art learning approaches which verifies its superiority in improving SEE for UAV-IRS-assisted LAE wireless communications network.

[IR-17] Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval

链接: https://arxiv.org/abs/2607.24165
作者: Sungguk Cha,DongWook Kim,Mintae Kim,Youngsub Han,Byoung-Ki Jeon,Sangyeob Lee
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Finding a long document relevant to a multi-part request is not the same as establishing that it contains every requested piece of evidence. We study this gap for conjunctive document retrieval, where two or three explicit conditions must be supported on different pages of one document. We use n-Clue as a controlled measurement instrument: 1,000 queries over 2,021 documents pair all-condition golds with naturally occurring documents that satisfy only a subset, and a complete-first success requires a top-10 gold to precede every released subset qrel. Across 70 configurations, condition-wise decomposition improves two dense backbones by 6.8–7.3 points and lexical–visual fusion adds 8.7, while four generic rerankers all reduce Gold-NDCG; these directions replicate on a four-source stress set. Scaling one dense family from 0.6B to 8B changes complete-first success by 0.0 points. The strongest displayed hybrid illustrates the resulting gap: it finds a gold for 81.1% of queries but succeeds complete-first on only 35.8%, and the gap persists across condition count, target length, candidate density, query rendering, and the four-source stress set. Finally, page-aware visual systems surface stored support for every condition on only 5.1–5.3% of queries. These results identify condition coverage, rather than gold discovery alone, as the central bottleneck.

[IR-18] Domain-Specific Data Quality Analysis Using Technology-Independent Query Templates

链接: https://arxiv.org/abs/2607.24151
作者: Arno Kesper,Lukas Sebastian Hofmann,Markus Matoni,Gabriele Taentzer
类目: Databases (cs.DB); Information Retrieval (cs.IR)
备注: 38 pages, 19 figures, submitted to SoSym Journal

点击查看摘要

Abstract:In an increasingly data-driven world, effectively working with data depends heavily on its quality. Quality analysis is a central aspect of data quality management. As data quality is typically domain- and context-specific, the definition of quality requirements is primarily the responsibility of domain experts. However, domain experts often lack the query language expertise needed to implement quality analyses. Therefore, the process of defining quality analyses results in a resource-intensive workflow that requires the involvement of technical experts, effectively excluding domain experts from independently managing data quality. To address this challenge, we present the Quality Pattern Model framework (QPM), a model-driven approach to define templates for data quality analyses that are independent of specific database technologies and application domains. QPM can eliminate the need for deep technical expertise and prevent the need for defining quality analyses several times for different database technologies. We present a proof-of-concept implementation of this approach for three database technologies: XML, RDF, and Neo4j. We evaluate the expressiveness of our approach, its applicability in the cultural heritage domain, and its usability by domain experts. For this purpose, we conducted a qualitative user study and empirically collected quality problems in a catalog. Our findings suggest that QPM matches and even exceeds the expressiveness of common database query languages. Furthermore, the results indicate that our tool enables domain experts to define template-based quality analyses independently, without requiring support of IT experts.

[IR-19] ConAlign: Conditional Alignment Framework for Balancing Biased and Unbiased Recommendation

链接: https://arxiv.org/abs/2607.24092
作者: Jingcheng Zhang,Yihan Wang,Qi Song,Liyin Hong
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Industry recommender systems trained on observational data suffer from various biases that create filter bubbles, causing user interests to collapse into narrow categories and severely degrading long-term engagement. While utilizing unbiased uniform data for debiasing has shown promise, existing methods remain impractical for industrial deployment due to limitations such as neglect of factual (biased) recommendation performance and the substantial computational overhead. To overcome these limitations, we propose ConAlign (Conditional Alignment Framework), a conditional debiasing approach for industrial deployment. The key innovation of ConAlign lies in a discrete gating-based conditional alignment mechanism that selectively transfers knowledge from the biased tower to the unbiased tower. Following a selective intervention paradigm rather than universal correction, it seamlessly balances factual accuracy and unbiased preference estimation while supporting real-time streaming adaptation. To the best of our knowledge, ConAlign is the first streaming debiasing recommendation framework successfully deployed in a large-scale industrial recommendation system that utilizes a small fraction of unbiased random traffic for debiasing. Extensive offline experiments on three real-world datasets rigorously validate the effectiveness of our proposed framework. Furthermore, large-scale online A/B testing on Kuaishou demonstrates significant improvements in long-term user engagement and interest diversity, with negligible latency overhead.

[IR-20] SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation

链接: https://arxiv.org/abs/2607.24025
作者: Yu Cui,Yi Xu,Jiahao Wang,Hao Zhang,Yu Zhang,Xiaoyi Zeng,Can Wang,Jinxin Hu,Jiawei Chen
类目: Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: 12 pages,7 figures

点击查看摘要

Abstract:Transformer architectures have achieved remarkable success across diverse domains; however, directly applying their standard self-attention mechanism to recommendation often yields suboptimal performance, sometimes even trailing behind well-designed simple recommendation models. In this paper, we reveal that this performance bottleneck stems from severe embedding and attention collapse unique to recommendation scenarios. The heterogeneity and long-tail nature of recommendation data lead to a severe spectral collapse dominated by a few principal singular values. We further theoretically demonstrate that this triggers a vicious cycle in recommendation model’s forward and backward propagation, which accelerates embedding and attention collapse and limits the model’s scaling capability with increased depth. To address these issues, we propose SpecFormer, a novel Spectral-Aware Transformer designed for mitigating embedding and attention collapse in recommendation. Specifically, SpecFormer introduces 1) a Learnable Spectral Softening module to dynamically smooth the singular values distribution of the input token embeddings; 2) a Spectrum-softened Attention mechanism to model feature interaction under a more uniform spectral distribution space; 3) a Spectral Residual Position Encoding via Taylor expansion of singular values, explicitly providing a spectral inductive bias for feature interactions. Extensive experiments on one industrial and two public datasets demonstrate that SpecFormer significantly outperforms state-of-the-art baselines. Notably, SpecFormer has been successfully deployed in a real-world commercial recommender system and exhibits exceptional scaling capabilities: stacking SpecFormer layers actively improves the attention effective rank and recommendation performance.

[IR-21] Mosaic: A Fleet of User Embedding Specialists for Recommendation at Meta RECSYS’26

链接: https://arxiv.org/abs/2607.24015
作者: John Zhiyuan Zheng,Xian Sun,Xiangyang Mou,Yujunrong Ma,Christina You,Michael Jiayuan He,Hrishikesh Paranjape,Aakarsha Agarwal,Hong Li
类目: Information Retrieval (cs.IR)
备注: Accepted in the 20th ACM Conference on Recommender Systems (RecSys '26),

点击查看摘要

Abstract:User representation is one of the highest-leverage modeling problems in industrial recommendation systems: a single advancement in how users are encoded can propagate across retrieval, ranking, and integrity tasks at platform scale. Prior industrial user representation work builds either a single user model that emits one or more embedding vectors or a shared backbone with task-specific adaptation. In this paper, we present Mosaic, a foundational user modeling platform that employs a fleet of specialists to learn user embeddings. The fleet comprises four architecturally diverse model families - memorization-driven, dense-heavy, sequential-based, and CoTrain models - each focusing on a distinct facet of user behavior. We developed MRM (Multi-task Relations Mining) and CRL (Cosine Redundancy Loss) techniques to maximize the marginal information contribution of each new specialist. We also introduce CoEval and User Tower Zero-Out, new logging-free embedding evaluation framework that improves development velocity while preserving downstream-aligned accuracy. Our hybrid CPU/GPU, online-and-offline serving stack allows each specialist to choose the adequate serving strategy to meet the freshness, latency, and computational requirements. Mosaic delivers consistent and significant offline NE improvements in addition to online gains.

[IR-22] MEMOIR: Temporal Behavioral Memory for Recommendation Across the Preference-Drift Spectrum

链接: https://arxiv.org/abs/2607.23986
作者: Younggue Bae
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL)
备注: 6 pages, 2 figures

点击查看摘要

Abstract:We propose MEMOIR, a framework that segments user interaction histories into temporal windows, generates semantic behavioral memory for each period using an LLM, and aggregates current state, evolution direction, and predicted future into a single user representation. On the Electronics and Clothing_Shoes_and_Jewelry categories of Amazon Reviews 2023, MEMOIR is statistically tied with UniSRec, the strongest baseline, on aggregate NDCG@10 (0.0643 vs. 0.0641), splitting the four reported metrics 2-2: MEMOIR leads NDCG@10 and MRR, UniSRec leads HR@10 and HR@20. An ablation study finds that no single architectural component - the evolution-preserving contrastive loss, its directional-consistency term, or temporal window segmentation itself - individually explains much of MEMOIR’s approximately 18% relative gain over ID-based SASRec; all four ablations land within 2% of the full model on aggregate NDCG@10. Stratifying test performance by a composite preference-drift score instead reveals where the gain concentrates: MEMOIR leads on ranking-quality metrics (NDCG@10, MRR) specifically among users at the high- and low-drift extremes of the distribution, while UniSRec leads the volume-oriented HR@10/HR@20 metrics across all drift strata and edges out MEMOIR on ranking quality in the middle band. We report this drift-stratified pattern, rather than the near-tied aggregate numbers or any single ablated component, as MEMOIR’s most substantive and reproducible finding, and surface why it holds as an open question for future work.

[IR-23] Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models and a Roster Instrument Captures 0.5% of It

链接: https://arxiv.org/abs/2607.23893
作者: Dmitrij Żatuchin(1,2) ((1) Department of Information Technologies, EUAS, Tallinn, (2) a href=“http://Rankfor.AI” rel=“external noopener nofollow” class="link-external link-http"this http URL/a, Tallinn)
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL); Computers and Society (cs.CY)
备注: 28 pages, 5 figures. Data: this https URL

点击查看摘要

Abstract:Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the question one level down, in categories where the buyer picks a person. It issued 2,400 grounded API calls in one two-hour window on 24 July 2026: 120 buyer-intent prompts, four models (GPT-5.6 Sol, Gemini 3.6 Flash, Perplexity Sonar Pro, Grok 4.5), five iterations each, four European markets and five query languages. Every response was coded for whether it named an individual professional, by a rule cascade that never consults a roster and that drops detections resolving to a same-named American city (precision 96.9%, recall 61.7%, so every rate below is a lower bound). All inference corrects for clustering within prompt: intraclass correlation 0.258, effective n 407 against a nominal 2,400. Models named an individual in 25.8% of responses. Category dominates: real estate 35.4% and car dealerships 32.9% against insurance 9.1% (chi-square 159.3, p = 5.8e-8 after correction). Models differ four-fold, from Grok 38.0% to Gemini 9.3%. Citation type predicts naming and citation volume does not: naming responses cite the individual’s own site 2.6 points more often (95% CI +1.4 to +3.9) and category portals 4.3 points more often, and cite firm-owned pages at the same rate (44.1% against 45.5%). On nine matched translation pairs, English prompts named an individual in 36.7% of responses against 15.6% for the same question in the local language (OR 3.14, clustered p = 0.074, so the direction is clear and the design cannot close it). A 939-person roster built from public LinkedIn search matched 128 of 27,293 name-shaped mentions (0.47%), 26 of the 939 people were ever named, and the roster-derived rates of 0.0% to 25.4% measure that overlap. Roster-based measurement of individual AI visibility sees a small and unrepresentative slice of what models do.

[IR-24] A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy 0 Tokens Bit-Exact Forever

链接: https://arxiv.org/abs/2607.23806
作者: Sietse Schelpe(Corbenic AI)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG); Performance (cs.PF)
备注: Industry experience report. 14 pages, 8 figures. Public testbench: this https URL ; companion repository with SHA-256 provenance manifest: this https URL

点击查看摘要

Abstract:Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is solved and has passed an independent verification step that never consults the answer key, every new instance of that family is answered at zero generation tokens, bit-exact, deterministically. Across 180 fresh instances spanning nine problem families, four architectures from four vendors - dense and mixture-of-experts - each score 180/180 at zero generation tokens per answer: execution-bound capability decoupled from parameter scaling. A negative control attributes the capability fully to the memory: emptied, it solves nothing. The same verify-before-store contract holds for open-ended reasoning: 88/88 consistency-gated acceptances across all four models, machine-checked formal proof, and reasoning-method transfer at 77/80. Memory selection takes 1.4 microseconds; a full reuse completes in 6-23 ms at 36 mWh. Approximate similarity retrieval selects the wrong item 94.3% of the time on a 4,500-item verified store where exact addressing makes zero errors. The store also serves as working context at a scale no shipped engine matches: a 6,000,000-token movable window on a single 46 GB GPU at flat memory, where vLLM stops at 30,399 tokens and SGLang silently truncates past 32,000. On published benchmarks, frontier models remain far ahead of any 12B at raw from-scratch reasoning; on everything this system has solved and verified, the comparison inverts: a frontier API call pays a fresh generation pass on every query, forever, while verified reuse costs zero tokens and returns the identical bits every time. A public testbench with free, rate-limited access accompanies this report: this https URL

[IR-25] ClawRec: A Claw-Native Recommender System

链接: https://arxiv.org/abs/2607.23779
作者: Chenghao Wu,Kesha Ou,Xiaolei Wang,Bowen Zheng,Bingqian Li,Enze Liu,Wayne Xin Zhao,Weitao Li,Long Zhang,Sheng Chen,Ji-Rong Wen
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Recommender systems have become integral to navigating the modern digital ecosystem. Yet most deployed systems remain confined within single-platform boundaries, observing localized interaction traces and ranking items from isolated candidate spaces. This design is poorly suited to real-world tasks that unfold through searches, content consumption, and comparisons across multiple information sources. Claw-style personal agents, with persistent access to authorized cross-platform context, create an opportunity for recommendation to operate around the user rather than any single platform. In this paper, we introduce Claw-native recommender systems, a new paradigm that moves beyond platform-local ranking to produce unified, complementary recommendation slates spanning diverse sources and content forms. To instantiate this paradigm, we present ClawRec, the first recommender system designed to operate natively in this environment. ClawRec maintains an evidence-linked, temporally structured user state that connects cross-platform behaviors with cross-source recommendations. It organizes retrieval around functional source roles and selects candidates according to their marginal utility, producing non-redundant slates aligned with the user’s active task. To enable rigorous evaluation, we introduce ClawRec-SimBench, a benchmark constructed from sequences of concrete life events and cross-platform behavior trajectories. Experiments show that ClawRec outperforms the strongest baselines, achieving an NDCG@20 of 0.6134 (+0.1126) and a Hit@20 of 0.6944 (+0.0854), while also improving user state quality and temporal alignment. Our code and dataset are available at this https URL.

[IR-26] Escaping the Euclidean Void: Manifold-Informed Flow Matching for Sequential Recommendation

链接: https://arxiv.org/abs/2607.23762
作者: Dengzhao Fang,Jingtong Gao,Yu Li,Xiangyu Zhao,Yi Chang
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Conventional recommenders capture users’ preferences by optimizing observed user-item relations, whereas continuous generative recommendation additionally learns the trajectory of synthesizing a target item. Flow matching drives this process by gradually shaping initial noise into a definitive next-item representation through intermediate states in a continuous embedding space. However, item catalogs are discrete and sparsely supported, meaning even a straight Euclidean path can cross continuous regions that contain little evidence of valid item semantics. Formalizing this failure as the Euclidean void, we propose MIRAGE, a Manifold-Informed Rectification framework for Accelerated Generation of Embeddings in sequential recommendation, which rectifies the learned embedding geometry around an unchanged straight probability path. By leveraging an item co-occurrence graph as a proxy for the underlying semantic manifold, MIRAGE aligns interpolated path states with local anchors, reorganizing the embedding space to ground the trajectory in valid item support. MIRAGE retains the original probability path and uses the graph only during training, thereby enabling accurate and efficient one-step inference. Extensive experiments on four real-world datasets reveal that MIRAGE consistently outperforms state-of-the-art baselines, effectively boosting performance on sparsely observed targets while achieving robust overall accuracy. Our code will be made publicly available upon publication.

[IR-27] Breaking the Loop: An Empirical Comparison of Strategies for Novelty and Freshness in YouTube Music

链接: https://arxiv.org/abs/2607.23749
作者: Srivaths Ranganathan,Zihuan Diao,Bernardo Cunha,Joshua L. Moore,Robin Dumas,Murat Goksedef,Yanwei Song,Mukai Lu,Gergo Varady,Tracy Pesin
类目: Information Retrieval (cs.IR)
备注: Accepted at 20th ACM Conference on Recommender Systems, September 27-October 02, 2026, Minneapolis, MN, USA. 8 pages

点击查看摘要

Abstract:Continuously trained ranking models in music recommenders fall into feedback loops where previously consumed items dominate recommendations. This suppresses two distinct content classes: new releases (temporal freshness) and unlistened catalog items (novelty). Industry practitioners have a wide menu of interventions available, ranging from serving-time heuristics, training-data reweighting, architectural debiasing, to uncertainty-driven exploration, each of which are well understood in academic settings. But live systems offer challenges with continuously ingested content, interconnected components, and practical limitations that counteract the findings from academic research. We report results from off-policy online A/B tests for six interventions and a combination experiment across four conceptual layers (serving, training, architecture, exploration) on the YouTube Music homepage. All interventions modify the ranking model or the serving layer that consumes its scores; candidate generation and other upstream components are held fixed. We discuss key takeaways from our results: first, serving-time interventions on continuously trained systems are neutralized by the learning loop. Second, architectural debiasing reduces popularity dominance and improves diversity but does not create discovery, while carrying hidden integration costs. Finally, uncertainty-driven exploration interventions with a Spectral-normalized Neural Gaussian Process (SNGP) head produce the largest new-release lift, though they come with a measurable engagement or diversity tradeoff. We close with recommendations on which layer to intervene at, and the hidden costs of each choice. Comments: Accepted at 20th ACM Conference on Recommender Systems, September 27-October 02, 2026, Minneapolis, MN, USA. 8 pages Subjects: Information Retrieval (cs.IR) Cite as: arXiv:2607.23749 [cs.IR] (or arXiv:2607.23749v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2607.23749 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Related DOI: https://doi.org/10.1145/3773078.3831876 Focus to learn more DOI(s) linking to related resources

[IR-28] Melo: A Production LLM -Powered Music Recommendation Agent

链接: https://arxiv.org/abs/2607.23718
作者: Shijia Wang,Da Guo,Qiang Xiao,Fanghui Bi,Weisheng Li,Dongjing Wang,Chuanjiang Luo
类目: Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:We describe Melo, an LLM-powered music recommendation agent deployed on NetEase Cloud Music. Melo is structured as a deterministic five-node state graph over heterogeneous tools, with a prompt- and state-machine-driven orchestration policy rather than a fine-tuned controller. At industrial scale, the bottleneck is not how smart the brain is but how the system detects and recovers from the mistakes that brain makes. Two production failure modes drove the design: entity hallucination, where the agent commits to interpretations unsupported by the live catalog or user-behavior index, and long-tail degradation, where over-constrained requests collapse to generic popular fallbacks. We address them with two complementary mechanisms. Inference-time entity grounding repurposes the production search index as a verification primitive that gates entity decisions before they propagate downstream. Reflective retry verbalizes failure reasons from a broken tool chain and feeds them into the next planning step, so the system can relax or revise constraints rather than fall back blindly. A one-month online A/B test across NetEase Cloud Music’s playlist surfaces reports an over 2 pp lift in a primary playlist retention metric and a lift of over one minute in a core playlist engagement metric. Offline ablation isolates a 7.8 pp reduction in entity misidentification from the three-layer grounding stack on our evaluation set, and a triggered-session analysis on our evaluation set shows reflective retry firing on 5.8% of sessions with 59% process-level recovery. Our deployment experience suggests that progress on LLM-powered music recommendation at this scale depends as much on the named, ablatable runtime machinery that catches and corrects the brain’s mistakes as on the brain itself: a hypothesis we offer for the community to test.

[IR-29] owards a Relevance Posterior in Neural Information Access SIGIR2026

链接: https://arxiv.org/abs/2607.23561
作者: Andrew Parry,Emmanouil Georgios Lionis,Debasis Ganguly,Sean MacAvaney
类目: Information Retrieval (cs.IR)
备注: SIGIR 2026 Perspectives Track

点击查看摘要

Abstract:Modern information retrieval systems typically operationalise relevance as a query-conditional score computed at inference time. This design choice has become dominant such that alternative decompositions of relevance are rarely discussed, despite the long history of document and query priors in probabilistic retrieval and large-scale search. As neural ranking models grow more computationally expensive and retrieval pipelines expand to include multi-stage ranking, recommendation, and retrieval-augmented generation, this monolithic view of query-time scoring becomes increasingly limiting. We argue that modern information access systems are more naturally understood as performing approximate posterior inference, in which relevance is refined through a staged combination of query-dependent likelihoods and query-independent priors. We extend classical probabilistic retrieval formalisms to contemporary learned systems and show how explicit likelihood-prior decomposition exposes new opportunities to shift computation offline while disentangling document-level and interaction-level beliefs. We present empirical evidence that incorporating query-independent document utility can complement existing rankers and improve effectiveness with minimal query-time computation (solely score fusion). Concretely, a learned prior improves first-stage retrieval through rank fusion (up to 0.046 nDCG@10 on TREC DL-2019 and 0.029 nDCG@10 on TREC DL-2020) and also improves downstream re-ranking, with the largest gains observed for the LLM re-ranker RankZephyr (up to 0.054 nDCG@10 on TREC DL-2020). Finally, we discuss how this decomposition connects to broader information access and outline research directions for designing retrieval systems that explicitly allocate modelling capacity between offline priors and online interaction.

[IR-30] Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework

链接: https://arxiv.org/abs/2607.23507
作者: Madhav S Baidya
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注: 33 pages, 1 figure, 20 tables. Technical report

点击查看摘要

Abstract:Choosing the right text embedding model is one of the most consequential – and most frequently under-examined – decisions in building a retrieval or search system, yet the model that tops a leaderboard is rarely the best choice for a given deployment. This report develops a practical, evidence-based framework for embedding model selection, built on a benchmarking study that evaluates T3EM (Text 3 Embedding Model), a commercial API-based embedding model, against a broad set of open-source alternatives on English-language retrieval tasks, and situates these findings within the wider Massive Text Embedding Benchmark (MTEB) landscape spanning classification, clustering, semantic similarity, reranking, pair classification, bitext mining, and summarization. Beyond raw benchmark scores, the report traces the full path from embedding model to retrieved result – how embeddings are produced, how they are indexed and searched at scale, and how document chunking strategy shapes retrieval quality – so that model choice can be reasoned about as one decision within a complete retrieval pipeline rather than in isolation. The result is a consolidated set of practical recommendations for selecting an embedding model according to task, latency, cost, and deployment constraints.

[IR-31] A Novel Gravity-Quasi-Laplacian Approach to Identifying Influential Nodes in Complex Networks

链接: https://arxiv.org/abs/2607.23419
作者: Shima Esfandiari,Seyed Mostafa Fakhrahmad
类目: ocial and Information Networks (cs.SI); Information Retrieval (cs.IR)
备注: 20 pages

点击查看摘要

Abstract:Identifying influential nodes in complex networks is a fundamental challenge with broad applications in areas such as social network analysis, communication infrastructure, transportation systems, and information networks. Existing ranking methods typically rely on combinations of structural features-such as degree, k-shell index, and neighborhood connectivity-to estimate a node’s importance. However, many of these approaches suffer from key limitations, including insufficient accuracy, low resolution in distinguishing nodes with similar influence, dependence on tunable parameters, and high computational complexity, which restrict their practicality in large-scale or real-world networks. This study introduces a new ranking framework that integrates a quasi-Laplacian structural measure with a gravity-inspired aggregation process. The core idea is to construct a strengthened representation of each node’s structural role using only simple yet informative attributes-namely degree and k-shell index-and then evaluate its local influence through a short-range interaction mechanism. The proposed approach is designed to be free of tunable parameters, interpretable, and computationally efficient, requiring only a small fixed gravity radius (R=3), which makes it suitable for large and diverse networks. Experiments conducted on nine real-world networks and compared against eight state-of-the-art methods demonstrate that the proposed framework consistently outperforms existing techniques in terms of accuracy, resolution, and computational simplicity. These results highlight the effectiveness of the gravity-quasi-Laplacian paradigm as a reliable and scalable tool for identifying influential nodes in complex networks.

[IR-32] Statistically Supported LLM Ingredient and Recipe Data Collection in Computational Nutrition

链接: https://arxiv.org/abs/2607.23273
作者: James Izzard,Hassan Eshkiki,Fabio Caraffini
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)
备注:

点击查看摘要

Abstract:Computational nutrition needs precise ingredient data, but current databases are incomplete, inconsistent, and built for human reference rather than automated reasoning. LLMs could help fill these gaps, but single-pass outputs are unreliable and can introduce silent errors into downstream computation. We present a quality-controlled LLM pipeline for ingredient data acquisition that combines robust statistical estimation, domain-specific invariant checks, and a web-fetch fallback. An illustrative Heap’s Law fit to 233 recipes suggests that unique-ingredient growth is sub-linear and front-loaded: the projected ratio of unique ingredients to recipes falls from 1.74 at 100 recipes to 0.19 at 5,000. For each ingredient attribute, repeated LLM queries are treated as samples from a model-induced answer distribution, and we apply robust point estimators and normalised confidence scores across numerical, Boolean, multiple-choice, open categorical, and optional integer types. An invariant guard layer enforces nutritional and logical self-consistency within each ingredient record. Minor numeric inconsistencies are reconciled via a linear program that minimises worst-case percentage deviation while preserving semantic zeros, and major violations are escalated to web-evidence-grounded repair, then human review only if that fails. On a curated 30-ingredient reference set, the pipeline achieves 98.4% exact match on nutrient flags and cuts median absolute percentage error on nutrient ratios from 31.9% for the median-aggregated baseline to 10.1%, a reduction of 21.8 percentage points, at an API cost of about 1 per ingredient. This frames LLM-assisted database construction as a controlled data-engineering workflow that makes uncertainty operational rather than discarding it.

[IR-33] SMART: LLM -Augmented Hybrid Retrieval for Dynamic Product Ads RECSYS’26

链接: https://arxiv.org/abs/2607.23121
作者: Congfei Zhang,Jingxiao Ma,Xiaodong Liu,Hsiang-wei Chao,Siman Wang,Ge Liu,Shantanu Aggarwal,Vincent Zhang,Meghana Missula,Rachel Liao,Zichu Li,Xiao Bai,Yunzhi Zhou,Yajun Wang,Zhe Liu,Jinchao Li,Yu Zhang
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: To be published in the 20th ACM Conference on Recommender Systems (recsys’26), September 27 - October 2, 2026, Minneapolis, MN, USA

点击查看摘要

Abstract:Dynamic Product Ads (DPA) require retrieving relevant items from multi-million product catalogs, balancing two competing objectives: retargeting (re-surfacing known interests) and prospecting (discovering new categories). While Large Language Models (LLMs) capture semantic intent better than traditional embedding models, deploying them at scale introduces prohibitive inference costs and lexical mismatch issues. Through controlled experiments on millions of users, we demonstrate a critical retrieval decomposition: rule-generated queries excel at retargeting on a lexical BM25 index, while LLM-generated queries excel at prospecting on a dense ANN index. Building on this, we propose SMART (SeMantic-aware Adaptive ReTrieval). To manage costs, a lightweight quality gate identifies coverage gaps in initial keyword results, adaptively routing only the ~10% of users who benefit from semantic prospecting to the LLM path. Offline evaluation demonstrates that this gated approach captures the bulk of semantic prospecting gains in Relevance Score while maintaining competitive re-targeting performance at a 90% reduction in LLM costs. Finally, in a 2-week online A/B test at Snap, SMART improved the ad conversion rate by +27.6% over a strong embedding-based baseline.

[IR-34] EGR: Embedding-Native Generative Retrieval with a Shared LLM RECSYS2026

链接: https://arxiv.org/abs/2607.23038
作者: Xiaodong Liu,Congfei Zhang,Hsiang-wei Chao,Siman Wang,Tong Zhao,Xiao Bai,Vincent Zhang,Jingxiao Ma,Zhe Liu,Wenfeng Zhuo,Zichu Li,Jitin Krishnan,Yunzhi Zhou,Yajun Wang,Jinchao Li,Yu Zhang
类目: Information Retrieval (cs.IR)
备注: Accepted to RecSys 2026

点击查看摘要

Abstract:Generative retrieval is increasingly popular in large-scale recommendation and advertising systems, yet current methods introduce practical complications. Semantic-ID methods rely on quantization, mutable identifier vocabularies, and token-to-item grounding; embedding-based pipelines train the item encoder separately from the query generator, which limits user-item alignment. We propose EGR, an Embedding-Native Generative Retrieval framework for recommendation and advertising. EGR uses a single shared LLM to learn item representations from item metadata and user representations from interaction histories in one embedding space. Items are indexed directly as dense vectors, and user histories are encoded as dense retrieval queries. Joint contrastive training groups related items and aligns queries with their target items. We evaluate EGR on public benchmarks, industrial data, and live deployment. EGR outperforms published baselines on Amazon Reviews; on Snap DPA, it scales with data, handles cold-start items, and benefits from multimodal input. In production, EGR delivers a +2.91% conversion-rate lift, simplifying system design while improving retrieval quality and ad performance.

[IR-35] VecTree-RAG : An Agent ic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy

链接: https://arxiv.org/abs/2607.23006
作者: Xinyan Zhong,Yuwei Shi,Yuqi Wei,Chen Shen,Tianhang Zhou,Zhenghao Wu
类目: Information Retrieval (cs.IR); Materials Science (cond-mat.mtrl-sci); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the supporting evidence within those papers. Conventional retrieval-augmented generation typically addresses both through similarity search over fixed-length passages, flattening document structure and separating scientific claims from their methodological and argumentative context. We present VecTree-RAG, an agentic framework that assigns these tasks to complementary retrieval mechanisms. Vector search ranks compact document and section representations across the corpus, whereas reasoning-guided traversal of source-verified section trees localizes evidence within shortlisted papers. Full text is retained in a page store and exposed progressively only after structural localization. We evaluate VecTree-RAG on 300 QASPER questions, an open-access subset of 54 LitQA2 questions, and 49 multi-document MOSAIC questions. Compared with Dense RAG, reranked Dense RAG, RAPTOR, and Search-o1, VecTree-RAG obtained the highest observed answer score on all three benchmarks, reaching 0.800 LLM-judge correctness on QASPER, 0.925 accuracy on LitQA2, and a 0.547 composite score on MOSAIC. On QASPER, its evidence-page precision was 0.274, compared with 0.046–0.071 for the baselines. LitQA2 ablations further showed that the complete vector–tree architecture required fewer inference tokens than variants without tree navigation or corpus-level vector routing. These results indicate that vector retrieval narrows the corpus-level search space and tree navigation concentrates reading on structurally relevant evidence. Although multi-turn inference remains more expensive than single-call retrieval, VecTree-RAG provides a structure-aware and traceable architecture for scientific literature question answering.

[IR-36] When Does Few-Shot Prompting Help? A Systematic Empirical Study of Shot-Count Effects Across Model Scale Architecture and Output Parsing Robustness

链接: https://arxiv.org/abs/2607.22969
作者: Ayush Dwivedi,Ashvi Soni
类目: Information Retrieval (cs.IR)
备注: 12 pages, 4 figures, 8 tables. Code available at this https URL

点击查看摘要

Abstract:Few-shot prompting, the practice of prepending a small number of input-output demonstration pairs to a query before presenting it to a large language model (LLM), is among the most widely adopted inference-time techniques in NLP. Yet little systematic work investigates how shot count interacts with model scale, architecture, and output format compliance in determining classification performance. This paper presents a controlled study of five LLMs across six shot-count configurations (k in 0,1,2,3,5,8) on the AG News four-class benchmark (n=200). Our models span proprietary and open-source families: Gemini Flash Lite, GPT-4o-mini, Llama 3.1 8B, Llama 3.3 70B, and Llama 4 Scout 17B. We report macro-averaged F1 with 95% bootstrap confidence intervals (B=10,000), permutation-test p-values, and Cohen’s d effect sizes across all 30 configurations. Our findings reveal four qualitatively distinct behavioral regimes: (1) models already well-calibrated at zero-shot that show modest, statistically insignificant gains (Gemini, GPT-4o-mini); (2) models that undergo catastrophic zero-shot failure but recover dramatically with a single example (Llama 3.1 8B, d=10.98, p0.0001); (3) models optimal at zero-shot that degrade monotonically with additional examples (Llama 4 Scout); and (4) models exhibiting a U-shaped curve (Llama 3.3 70B: 0-shot F1=0.907, 2-shot F1=0.635, 5-shot F1=0.785 with parser corrected). We additionally identify, diagnose, and correct a systematic parsing artifact that artificially deflated Llama 3.3 70B performance by up to 206%, constituting a methodological contribution to LLM evaluation practice. Our results demonstrate that the relationship between shot count and classification performance is not monotonic, not universal, and not predictable from model scale alone.

[IR-37] Beyond Exact Match: How Evaluation Methodology Dominates Model Choice in LLM -Based Product Attribute Extraction

链接: https://arxiv.org/abs/2607.22949
作者: Ayush Dwivedi,Ashvi Soni
类目: Information Retrieval (cs.IR)
备注: 9 pages, 3 figures, 8 tables. Preprint under review. Code available at this https URL

点击查看摘要

Abstract:Large language models (LLMs) have become a default choice for structured product attribute extraction in e-commerce pipelines, with practitioners reporting widely varying performance across models, datasets, and prompting strategies. This paper presents a controlled empirical study comparing four prompting strategies – zero-shot, few-shot, schema-guided, and definition-augmented – across two production-grade LLMs (GPT-4o-mini and Gemini 2.5 Flash) on the MAVE benchmark. We evaluate 6,400 attribute-level predictions using both exact and fuzzy string matching, and conduct a rigorous noise audit of the ground truth labels. We formally decompose F1 variance across four experimental factors and find that evaluation methodology produces variance approximately 23 times larger than model choice and 5 times larger than prompting strategy choice. We further establish that the MAVE benchmark exhibits a 23.2% ground truth noise rate against modern LLM outputs. Paired permutation tests (B=10,000) confirm that the inter-protocol F1 gap is highly significant (p0.0001) and Cohen’s kappa of 0.769 between protocols indicates substantial agreement. We conclude that for production attribute extraction pipelines, evaluation methodology and data quality dominate the impact of model selection and prompt engineering.

[IR-38] Language-Routed RAG and Direct Option Scoring for Multilingual Financial QA: DS@GT at FinMMEval

链接: https://arxiv.org/abs/2607.22841
作者: Justice Ayela,Kabir Sahni
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Accepted for publication in the CLEF 2026 Working Notes. 15 pages, 3 figures

点击查看摘要

Abstract:We present DS@GT’s submission to FinMMEval 2026 Task 1, a multilingual financial exam question answering benchmark spanning English, Spanish, Greek, Chinese, and Hindi. Financial certification exams such as the CFA, EFPA, and CPA demand structured domain reasoning that standard NLP benchmarks do not capture, and this challenge compounds across languages where retrieval and representation infrastructure is underdeveloped. We build a retrieval-augmented pipeline on LangGraph that detects query language and retrieves semantically relevant exemplars from a 30,209-entry multilingual knowledge base using BGE-M3 embeddings and FAISS indexing. The system then scores answers via Retrieval-Augmented Direct Scoring (RADS), reading next-token log-probabilities over candidate option letters rather than generating free-form output. For low-resource languages, we fuse per-language and cross-lingual retrieval indices using weighted Reciprocal Rank Fusion. Model selection is language-routed: Qwen3-14B for Arabic, Chinese, and Hindi; Qwen2.5-14B for English; and Llama-3.1-8B for Greek, a routing derived from empirical ablations that reveal substantial language-asymmetric performance gaps. Notably, chain-of-thought prompting significantly degrades Greek accuracy (90.7% to 20.9%), and enabling Qwen3’s default thinking mode collapses Arabic RADS performance to near-chance levels. Our results indicate that effective multilingual financial reasoning requires language-aware retrieval, model routing, and deliberate scoring strategy selection.

[IR-39] MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation CIKM

链接: https://arxiv.org/abs/2607.22706
作者: Hyewon Lee,Minkyung Song,Junghyun Oh,Seunghoon Han,Sungsu Lim
类目: Artificial Intelligence (cs.AI); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
备注: 12 pages, 1 figure, 7 tables. The 1st International Workshop on Retrieval-Driven Generative AI ScienceON AI Challenge 2025@CIKM

点击查看摘要

Abstract:This paper presents the MPR-CiteG framework, which achieved second place in the ScienceON AI Challenge by addressing two fundamental challenges in generative AI: inefficient retrieval and the absence of source verification. We propose a dual-component system, termed MPR-CiteG, in which the Multi-Portfolio Retriever (MPR) efficiently retrieves diverse and relevant information, while the Citation-Grounded Generation (CiteG) module ensures that every generated output remains factually consistent and explicitly attributed to its source. MPR-CiteG represents a significant step toward building more trustworthy and accurate LLMs that are not only capable of generating information but also of grounding their responses in reliable evidence, thereby mitigating common issues like model hallucination. Extensive experiments on the challenge dataset validate the effectiveness and reliability of our approach. Our code is available at this https URL.

[IR-40] owards Nexus-Score: Metadata Gaps Limit Scholarly AI Attribution

链接: https://arxiv.org/abs/2607.22684
作者: Aadi Narayana Varma Dantuluri,Sushrut Thorat,Paras Chopra
类目: Digital Libraries (cs.DL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注: 19 pages, 4 figures, 10 tables; includes supplementary methods and reproducibility information

点击查看摘要

Abstract:Artificial intelligence systems increasingly mediate how science is found and credited. We asked whether missing metadata prevents AI systems from crediting work. As a boundary test, an AI system citing without access to task-relevant paper lists often produced out-of-list identifiers, some fabricated. We then tested the mechanism in real scholarly infrastructure by using OpenAlex records to hide or restore author, institution, funder, reference, and text-access links while holding works and tasks fixed. Restoring the relevant link made the corresponding attribution possible; restoring the wrong kind did not, with 0 correct answers across 469 completed mismatched tests. Thus, in these tasks, one metadata facet did not substitute for another. Missing links led to invented answers, refusals, or tool-budget exhaustion, and web search did not recover hidden author links. In sum, AI systems credited work only when record connections were visible or recoverable. This motivates Nexus-Score, a record-level check for metadata gaps, to guide repair and help prepare the scholarly record for AI-mediated use.

[IR-41] Structure Over Scale: Schema-Constrained Causal Graphs for RAG

链接: https://arxiv.org/abs/2607.22592
作者: Marc Saouda,Rajprakash Bale,Eren Aldis,Cloves Almeida(Boston Consulting Group)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注: 26 pages, 6 figures

点击查看摘要

Abstract:Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and relationships exhaustively, producing graphs whose size and construction cost scale with corpus length rather than with the reasoning a query requires. We introduce HCG-RAG (Hierarchical Causal Graph RAG), which replaces open-ended extraction with schema-constrained causal graphs: an automated pipeline distills a corpus into a fixed, typed vocabulary of causal variables and materializes a compact two-tier graph over it. Our schema-constrained graphs match entity-relation baselines on answer quality at a fraction of the cost: 3-20x fewer nodes, 8x-135x fewer build-time LLM calls than the most LLM-intensive baseline (MS-GraphRAG), and graphs compact enough for a domain expert to audit, correct, and extend. On medical and clinical benchmarks, including a neurologist-validated epilepsy dataset, HCG-RAG matches or exceeds the best entity-relation systems. An ablation isolates the causal graph as a structured retrieval filter, contributing +6 percentage points (pp) over embedding-only retrieval. Across all domains with discoverable hierarchical causal structure, only methods imposing higher-level organization outperform flat entity-relation retrieval, indicating that what is placed in the graph matters more than how many nodes it contains.

[IR-42] oo much evidence too little time: From text to actionable recommendations through multi-objective evidence reasoning

链接: https://arxiv.org/abs/2607.22574
作者: Adela Bara,Simona-Vasilica Oprea
类目: Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Evidence-based clinical decision making requires specialists to identify, evaluate and synthesize relevant scientific literature. However, PubMed searches for complex clinical cases often return hundreds of publications that cannot be reviewed manually under time constraints. This study proposes SCEPTER (Single-Case Evidence-driven PubMed-To-rEcommendation Reasoner), a framework for transforming clinical case descriptions into evidence-based recommendations. SCEPTER combines PubMed retrieval, PubMedBERT semantic ranking, large language model (LLM)-based claim extraction, evidence-level weighting, contradiction detection, consensus analysis and multi-objective Pareto claim selection. The framework generates structured evidence syntheses and grounded actionable recommendations. A Paper QA module further enables interactive exploration of selected publications. The proposed framework introduces multi-objective reasoning model that integrates literature support, contradiction analysis and interactive literature interrogation into a unified clinical decision-support pipeline. Evaluation on 150 case studies demonstrated that the framework reduced an average search space of 576 papers to 53 retained papers, 7 Pareto-optimal claims and 3 final recommendations, corresponding to an overall compression ratio of 192:1. Despite this reduction, the retained evidence maintained high diversity (entropy=0.901). The ablation study showed that Pareto-based selection increased evidence diversity and recommendation utility compared with conventional ranking approaches.

[IR-43] A corrective agent ic hybrid RAG and an operations-grounded evaluation for a scientific facility

链接: https://arxiv.org/abs/2607.24663
作者: Rajat Sainju,Dariusz Jarosz,Hairong Shang,Michael Prince,Ryan M. Aydelott,Mathew J. Cherukara,Yine Sun,Michael D. Borland
类目: Accelerator Physics (physics.acc-ph); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We present APS-RAG, Advanced Photon Source Retrieval Augmented Generation, a deployed platform that makes the institutional knowledge at the Advanced Photon Source (APS) accessible to staff through natural-language queries, along with an operations-grounded evaluation. The retrieval engine fuses dense, sparse, and knowledge-graph (KG) channels with query-type-adaptive reciprocal-rank fusion, adds a corrective agentic loop, and runs a native-tool ReAct executor over a Model Context Protocol (MCP) tooling layer. We construct APS-Bench, a 50-question, question-answering (QA) dataset with auditable gold answers. Every retrieval-augmented variant numerically improves strict vital-nugget recall over a naive BM25 baseline (63.8%), with the full corrective Agentic GraphRAG scoring (70.3%). The cross-encoder reranker contributes significantly to answer quality: removing it and allowing the LLM to score relevance drastically reduces strict vital recall by 32.8%. The graph channel and corrective loop contribute positively as expected, but the performance gains are marginal. Additionally, we also compare the performance of open-source and closed-source LLMs in final answer synthesis. We release the APS-Bench construction methodology, the six-layer evaluation harness, and the underlying codebase, along with the ‘/aps-rag’ retrieval agent skill framework, to support reproduction and adoption at other facilities. Together, the deployed platform and its operations-grounded evaluation present a promising workflow for trustworthy, statistically grounded AI assistance in facility operations, transferable to other large scientific instruments.

[IR-44] Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature

链接: https://arxiv.org/abs/2607.23886
作者: Tanjin He,Aikaterini Vriza,Logan Ward,Xu Huang,Yiming Chen,Anubhav Jain,Gerbrand Ceder,Rajeev S. Assary,Ian T. Foster,Maria K. Y. Chan
类目: Materials Science (cond-mat.mtrl-sci); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
备注:

点击查看摘要

Abstract:X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through fragmented textual context in the literature. Here, we use multimodal (image and text) literature mining to transform this dispersed knowledge into an AI-ready experimental data resource. We developed a scalable spectroscopy data digitization pipeline that identifies XAS figures in full-text articles, digitizes spectral curves, and links each spectrum to accompanying metadata on the measured edge and material. Applying this pipeline to the battery literature produced an open dataset of 13,740 XAS spectra, spanning 66 absorbing elements and diverse battery chemistries, with expert validation confirming accurate extraction of spectral and metadata information. By converting literature-embedded spectra into structured numerical data, this dataset provides a foundation for large-scale XAS analysis, cross-laboratory comparison, high-throughput characterization, and autonomous discovery of advanced materials.

[IR-45] Route Based Map Matching via a Structured Codebook and Token Sequence Decoding

链接: https://arxiv.org/abs/2607.22543
作者: Takara Sakai
类目: Optimization and Control (math.OC); Information Retrieval (cs.IR); Networking and Internet Architecture (cs.NI)
备注: 18 pages, 9 figures

点击查看摘要

Abstract:This study proposes an efficient and computationally light route based map matching method for GPS track data on urban expressway networks. The key idea is to exploit a symbolic structure of named lines and named junctions that link level map matching leaves unused. We represent each candidate route as a sequence of line and junction names, take the set of such sequences as a route codebook, and formulate map matching as scored alignment of a probe trajectory against members of the codebook. Probes become token sequences via a mesh quantizer, a precomputed grid mapping each coordinate to a line or junction token, and the decoder returns a member of the codebook by construction. The codebook is indexed by a DAFSA \times Levenshtein automaton, a fuzzy lookup technique from approximate string matching and speech recognition; the per query decoding cost is orders of magnitude lower than a brute force scan. We evaluate the method on a deformed replica of the Tokyo Metropolitan Expressway topology. The method recovers the exact route at moderate GPS noise and continues to identify the line and junction sequence under heavy noise; a sensitivity analysis maps the mesh resolution operating range. Real probe evaluation, channel model calibration, and a head to head HMM comparison are left to a forthcoming version.

人机交互

[HC-0] Make or Take: How Students Navigate Self-Created and Instructor-Provided Cheat Sheets

链接: https://arxiv.org/abs/2607.24736
作者: Helen Weixu Chen,Victoria Sakhnini,Lesley Istead
类目: Human-Computer Interaction (cs.HC)
备注: This paper has been accepted for publication in ACM Transactions on Computing Education (TOCE)

点击查看摘要

Abstract:The use of cheat sheets in exams is often framed as a way to reduce cognitive load and support student performance. However, little is known about how students choose between self-created and instructor-provided cheat sheets, or how these choices relate to their broader approaches to exam preparation. We conducted a longitudinal study in a senior-level undergraduate software requirements course, where students could use either an instructor-provided or a self-created cheat sheet for both the midterm and final exams. Across three survey waves, we received 53, 50, and 44 responses, respectively. 41 students completed all three surveys and formed the longitudinal cohort used to examine how choices and experiences evolved over time, while exam-specific analyses used all available responses from the corresponding wave. Our findings identify several considerations that shaped students’ choices, including trust in instructor expertise, the desire for personalization, and preparation efficiency. We further show how students’ attitudes shifted over time and how their preferences were reflected in patterns of cheat sheet use, perceived content coverage, and challenges encountered during the exams.

[HC-1] Leveling the Playing Field: Temporal Video Segmentation for Individuals with ADHD in Computing Education

链接: https://arxiv.org/abs/2607.24612
作者: Veronica Pimenova,Chris Lee,Baramee Bhakdibhumi,Simon Chu,Andrew Begel
类目: Human-Computer Interaction (cs.HC)
备注: 16 pages, Accepted to 28th International ACM SIGACCESS Conference on Computers and Accessibility

点击查看摘要

Abstract:Individuals with Attention-Deficit/Hyperactivity Disorder (ADHD) often face significant barriers in computing education. In asynchronous learning environments, instructional videos can impose high extraneous cognitive load, often relying on assumptions about sustained attention and working memory that do not align with ADHD neurocognitive profiles. In this work, we evaluate a post-hoc video processing intervention that segments instructional content into single-instruction chunks followed by fixed-length pauses to reduce cognitive load. In a within-participants controlled study with 17 individuals with ADHD and 10 without, we find that the intervention has an equalizing effect. Although it improved performance for all participants, gains were larger for those with ADHD, reducing their errors and hesitations to levels comparable to those of participants without ADHD under the same intervention. These results align with the goals of Universal Design for Learning (UDL), by showing that cognitively-aligned, post-hoc instructional video modifications can reduce performance disparities across diverse neurocognitive profiles.

[HC-2] Characterizing In-the-Wild Personal Listening Device Use to Inform Earable Application Design

链接: https://arxiv.org/abs/2607.24603
作者: Supraja Ramesh,Jonas Hummel,Silvia Becker,Christopher Clarke,Jake Stuchbury-Wass,Kai Kunze,Axel Loewe,Michael Beigl,Tobias Röddiger
类目: Human-Computer Interaction (cs.HC)
备注: 20 pages with 2 pages appendix

点击查看摘要

Abstract:Ear-worn devices are evolving from audio-playback tools into sensing platforms for health, interaction, and context-awareness. Yet, earable systems are typically designed and evaluated under strong assumptions about how long, how often, and in which situations people actually wear personal listening devices (PLDs). To ground these assumptions in-the-wild behavior, we combine a survey of 330 adults with multi-year, passively logged headphone audio-exposure records donated via Apple Health by 90 of them. We characterize where and when people use PLDs, how logged use has changed in recent years, and how psychological traits and social context associate with PLD usage. Our results show that logged mean daily use has increased from 37 minutes in 2020 to 64 minutes in 2024. Listening was intermittent: no listening was logged on 48% of participant-days in 2024, and sessions were fewer but longer on weekends. Sensation seeking, particularly disinhibition, showed small to medium positive associations with self-reported PLD use. Younger adults listened at higher volumes than the 25-34 group. High-volume exposure was uncommon, with only 4% of participant-weeks exceeding World Health Organization (WHO) safe-listening limits. Finally, most participants also reported avoiding PLD use in social situations. We translate these findings into implications for earable computing: realistic expectations of intermittent rather than continuous wear, contextual coverage that anticipates systematic gaps, targeted safe-listening interventions, and personalization grounded in psychosocial and demographic profiles rather than assumptions of uniform use.

[HC-3] Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review ISSTA ISSTA093 ISSTA2026

链接: https://arxiv.org/abs/2607.24601
作者: Zhenhan Gao,Marvin Muñoz Barón,Umm-e Habiba,Daniel Graziotin,Stefan Wagner
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: 23 pages, 4 figures, 5 tables. To appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA093 (ISSTA 2026). Published under CC BY 4.0. Replication package: this https URL

点击查看摘要

Abstract:Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to understand. Developers struggle to assess the validity of LLM-generated reviews, making it difficult to gauge how much trust to place in them. The role of Explainable AI (XAI) in code review and its impact on trust remain underexplored. Objective: We study the influence of XAI on developer trust in AI-assisted code reviews. Method: We conducted a within-subjects user study with 34 participants, comparing three LLM-based code review systems with varying levels of XAI support: Condition A (detailed explanation and review feedback), Condition B (review feedback only), and Condition C (no explanations). Participants reviewed real-world code change requests alongside the AI-generated reviews. We measured trust perceptions, agreement with the AI recommendation, the reasoning given for each decision, and the time taken. Results: The level of explanation significantly influences both trust and agreement with AI recommendations, but in different ways. Full explanations (A) yield the highest perceived trust (M = 3.99/5) but not the highest agreement, whereas moderate explanations (B) achieve the highest agreement (89.22%). This could suggest that more explanation prompts developers to question AI recommendations more frequently. No explanations © results in the lowest trust and agreement. Explanation level did not significantly affect review time. The most commonly cited reasons for decisions were code readability and correctness. Conclusion: Incorporating XAI into code review significantly changes trust perceptions and agreement with AI recommendations. These results inform the design and evaluation of trustworthy AI-based code review systems, as well as studies on the human factors of AI-assisted software development. Comments: 23 pages, 4 figures, 5 tables. To appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA093 (ISSTA 2026). Published under CC BY 4.0. Replication package: this https URL Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC) ACMclasses: D.2.5; H.1.2; I.2.7 Cite as: arXiv:2607.24601 [cs.SE] (or arXiv:2607.24601v1 [cs.SE] for this version) https://doi.org/10.48550/arXiv.2607.24601 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Marvin Muñoz Barón [view email] [v1] Mon, 27 Jul 2026 16:04:24 UTC (900 KB) Full-text links: Access Paper: View a PDF of the paper titled Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review, by Zhenhan Gao and 4 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.SE prev | next new | recent | 2026-07 Change to browse by: cs cs.AI cs.HC References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[HC-4] Designing Within the Lines: Practitioners Perspectives and Visualisation Tool Evaluation in the Arabic Context IEEE-VIS2026

链接: https://arxiv.org/abs/2607.24571
作者: Muna Alebri,Noëlle Rakotondravony,Yassine Bechqito,Georgia Panagiotidou,Lane Harrison,Hassan Aldhanhabi,Salem Alkaabi
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注: Accepted for publication at the IEEE VIS 2026 conference

点击查看摘要

Abstract:Design guidelines and best practices serve as references that support designers throughout the visualisation design process. While considerable effort has identified the elements that contribute to effective data visualisations, little attention has been paid to how language (scripts and reading direction), tool support, and cultural context also shape design decisions. As a result, assumptions of homogeneity persist, with visualisation practices predominantly benefiting users of English and left-to-right (LTR) scripts while overlooking the needs of over two billion Arabic script users. We investigate how Arabic-speaking visualisation practitioners design for right-to-left (RTL) scripts. We report on an analytical evaluation of seven popular GUI-based visualisation authoring tools using an Arabic dataset complemented by interviews with 11 Arabic-speaking practitioners across journalism, design, and data analysis. Our findings reveal that visualisation practitioners constantly negotiate tensions between Arabic reading conventions, “universal” LTR visual norms, and limited tool support for Arabic text, Eastern numerals, and maps. They engage in substantial labour, such as manually mirroring charts, fixing alignment issues, and stitching together multi-tool workflows, while making strategic compromises in language choice, interactivity, and chart type. Our tool analysis further reveals fragmented, inconsistent support for RTL mirroring, poor numeral rendering, and map defaults that encode geopolitical assumptions. In light of these findings, we discuss how RTL visualisation work is carried out under many constraints that affect agency and creativity. We argue that visualisation tools and defaults operationalise linguistic and geopolitical power in RTL contexts, and offer research directions and design implications that more robustly support RTL practitioners.

[HC-5] Proceedings of The First Reflection in Creative Experience (RiCE) Workshop

链接: https://arxiv.org/abs/2607.24558
作者: Corey Ford,Olga Sutskova,Samuel Rhys Cox,Sarah Sterman,Max Kreminski,Rosa Van Koningsbruggen,Anqi Wang,Ege Otenen,Karly Ross,Giulia Di Fede,Yinmiao Li,Salvatore Andolina,Marit Bentvelzen,Pan Hui,Nick Bryan-Kinns
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Reflection and metacognition are central to the creative user experience. However, most HCI research on reflection focuses on clear, task-oriented goals such as to reflect on personal data or pedagogical outcomes. This contrasts with the open-ended and challenging to articulate goals of creative user experiences. For the first time, this workshop brings together interdisciplinary researchers, designers, educators, and artists across HCI, Cognitive Science, Design, AI, Learning Sciences, and Digital Art to examine reflection in creative interaction. The workshop will discuss themes, drawn from earlier discussions with HCI researchers and artists, on: how best to capture reflection in creative contexts, how to leverage the arts to support reflection for ethical change, and how to design creative AI that enhances - not hinders - critical thinking. By bringing interdisciplinary perspectives on reflection into discussion, the workshop will develop a guiding taxonomy for reflection in creative interaction to inform future creative practice and tool development.

[HC-6] RemiAssist: A Therapist-Supporting System for Photo-Based Reminiscence Therapy in Dementia Care

链接: https://arxiv.org/abs/2607.24536
作者: Shuchang Xu,Minglong Tang,Junyan Mao,Xiaofu Jin,Wazeer Zulfikar,Yasith Samaradivakara,Jiayi Zhou,Huamin Qu,Yuling Sun,Pattie Maes
类目: Human-Computer Interaction (cs.HC)
备注: Accepted to UIST 2026

点击查看摘要

Abstract:Despite growing interest in applying AI to photo-based reminiscence therapy (PRT) for people with dementia (PwD), existing systems primarily focus on PwD-AI interaction and often overlook therapists’ critical role in practical PRT delivery. We present RemiAssist, a system that supports therapist-in-the-loop PRT through AI-assisted planning and real-time facilitation. RemiAssist incorporates two core techniques: (1) a Memory Graph, which organizes key life events from a PwD’s photo collection into a hierarchical graph to support theme-centered intervention planning; and (2) a Context-Aware Guiding Strategy, which provides real-time suggestions to help therapists guide reminiscence conversations and respond to sensitive situations. A field study with eight therapist-PwD dyads suggests that RemiAssist was associated with a 44% improvement in planning efficiency and a 54% increase in conversation duration, and provided timely support for handling sensitive situations. We highlight opportunities for AI systems to empower therapists and enable more personalized reminiscence therapy in dementia care.

[HC-7] Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis ACM-MM2026

链接: https://arxiv.org/abs/2607.24430
作者: Yifan Hu,Shuwei He,Rui Liu,Haizhou Li
类目: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
备注: 10 pages, 5 figures, 5 tables. Accepted by ACM MM 2026

点击查看摘要

Abstract:Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, facial expressions encode subtle and rich affective cues that are crucial for empathetic speech interaction, whereas existing approaches often overlook this important modality. In addition, the lack of large-scale natural conversational datasets with both speech and visual modalities also limits the development of visual affect understanding in conversational this http URL address these limitations, we propose FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone. To efficiently encode facial expressions, we propose AUTokenizer, a single-codebook visual tokenizer that discretizes each frame-level facial expression into a compact token, trained with supervision from combinations of facial Action Units. We further introduce a dual direct preference optimization (DualDPO) strategy, which extends the DPO by jointly imposing preference constraints on both visual and speech token sequences, to enhance the model’s understanding of facial expressions and speech semantics in multimodal conversational contexts. Moreover, we construct VSDD-1K, a large-scale multimodal dialogue dataset collected through a fully automated pipeline from real-world Internet conversations, comprising over 1,033 hours of synchronized speaker videos and speech, with more than 85% of frames containing valid faces. Extensive objective and subjective experiments demonstrate that FacialTalker consistently outperforms strong baselines in facial-expression perception and speech synthesis quality, generating speech that is more natural, expressive, and better aligned with the conversational context. The results also validate the effectiveness of our training strategy and dataset construction pipeline.

[HC-8] Order-Bound Companionship: The Practice of Emotional Labor in Professional Game Companionship

链接: https://arxiv.org/abs/2607.24363
作者: Xiaohe Mo,Yujie Zhang,Yizhen Li,Jianyi Wang,Qinyi Liao,Ray LC
类目: Human-Computer Interaction (cs.HC)
备注: To appear in Proceedings of the ACM on Human-Computer Interaction (PACMHCI), Volume 10, Issue 7, Article GAMES050

点击查看摘要

Abstract:Labor in platform gig economy increasingly involves services involving relationship that demand significant emotional investment. Grounded in China’s unique socio-cultural and multi-platform context, this study explores professional game companionship, an under-explored digital labor practice. Through interviews with 22 game companionship practitioners, we used a micro-level perspective to relational gig work to analyze how workers navigate intimate boundaries and stakeholder networks. We found that companions adopt an “order-bound” mechanism: performing immersive deep acting during paid sessions, followed by complete emotional disengagement post-order. We also identified a tripartite companion-centric network featuring scenario-based performances with clients, competitive-symbiotic peer relations, and interdependent governance with companionship clubs. Furthermore, significant identity fluidity exists, with individuals frequently transitioning between companion, client, and club operator roles. We provide implications for future labor governance and platform designs for intimate digital work that explicitly account for institutionalized boundaries and gig workers’ shifting psychological needs.

[HC-9] Modeling Duelling Contagions of True and False Information in the Face of Inherent Individual biases

链接: https://arxiv.org/abs/2607.24360
作者: Vaibhav Krishna,Hirokazu Shirado,Feng Fu,Nicholas A. Christakis
类目: ocial and Information Networks (cs.SI); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Advanced digital communication has revolutionized how people create and consume information, making information diffusion an important topic of research for domains from public health to national security. Real-world scenarios of information diffusion often involve competing narratives - true and false - spreading simultaneously. We propose a novel agent-based co-diffusion model, grounded in “complex-contagion” and “spiral of silence” theories, to capture how network dynamics exploit cognitive biases to shape such interactions. Our findings reveal that manipulative narratives dominate when early spreaders hold them. These network dynamics further exploit inherent cognitive biases to amplify information diffusion regardless of veracity. Further, while favourable previous experience strengthen collective optimism, unfavourable experiences attenuate optimism only modestly. However, we found that early seeding of agents with lower self-censorship not only constrains the spread of manipulation but can also lead to dominance of well-informed populance. This has implications for policies that aim to facilitate healthier discourse, strengthen social cohesion, and ensure equitable access to reliable information.

[HC-10] Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim

链接: https://arxiv.org/abs/2607.24190
作者: Steve Aschenbrenner,Marcel Heisler,Thomas Sievers,Christian Becker-Asano
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: Acceoted at the 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026) Kitakyushu, Japan

点击查看摘要

Abstract:Social robots that rely on large language models for conversation are unable to retain information across sessions. This absence of memory violates social expectations, potentially preventing the formation of persistent relationships. This paper presents a lightweight episodic memory module that integrates vector-based semantic retrieval with an LLM-controlled dialog system, deployed on the humanoid robot head Kim. The module employs a hybrid scoring function combining cosine similarity with a memory strength metric to retrieve contextually relevant past interactions and inject them into the generation prompt. The system was evaluated in a within-subjects video-based online study (N = 43) using the Human-Robot Interaction Evaluation Scale (HRIES). Results show that episodic memory significantly increased perceived sociability (d = 0.60, p .001), with the strongest effects on perceived trustworthiness (d = 0.62) and warmth (d = 0.56). Perceived disturbance remained unchanged (d = 0.00), indicating that the implemented approach to personalized recall did not trigger privacy-related discomfort or uncanny valley effects. These findings suggest that episodic memory serves as a social lubricant in embodied Human-Robot Interaction, enhancing relational quality without eliciting negative affective responses.

[HC-11] EEGForceFusion: Joint Tokenised-Continuous Representation Learning for Subject-Independent Grasp Force Decoding

链接: https://arxiv.org/abs/2607.24126
作者: Sankalp Sunil Turankar,Yogesh Kumar Meena
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)
备注: 6 pages, 9 figures, 4 tables, accepted at Brain-Machine Interface (BMI) Systems Session, IEEE International Conference on Systems, Man, and Cybernetics (IEEE SMC 2026)

点击查看摘要

Abstract:Brain-machine interfaces provide a link between neural activity and external devices, enabling restoration of motor function and advancing human-machine interaction using non-invasive electroencephalography (EEG). However, continuous grasp force decoding remains challenging due to complex temporal dynamics, high inter-subject variability, and limited generalisation of existing approaches. To address this, we propose a hybrid EEG decoding framework that jointly models continuous and tokenised representations, enabling capture of both fine-grained neural structure and long-range temporal dependencies. The proposed approach integrates convolutional-recurrent representation learning, quantisation-based tokenisation, and transformer-based temporal modelling within a unified fusion-based regression architecture. Experimental evaluation on the WAY-EEG-GAL dataset under strict leave-one-subject-out conditions achieves R^2 = 0.817 in offline settings and R^2 = 0.793 in simulated real-time evaluation, with latency suitable for real-time deployment. These results demonstrate strong cross-subject generalisation and highlight the practicality of hybrid continuous-tokenised representations for real-time EEG-based force decoding in assistive robotics, neuro-rehabilitation, and human-machine interaction.

[HC-12] A Case Study on the Acceptance of a Humanoid Robotic Head Employed in Three Public Spaces

链接: https://arxiv.org/abs/2607.24113
作者: Marcel Heisler,Luca Randecker,Christian Becker-Asano
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: accepted at the 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026) Kitakyushu, Japan

点击查看摘要

Abstract:Previous research has shown that a human-like robot’s acceptance heavily depends on the setting in which it operates and its ability to perform relevant tasks. This paper, first, reports on how our robot processes natural language to generate a multimodal, verbal response integrating emotional expressions based on an emotion simulation backend. Then, it describes how visitors were invited to speak with our robot in their own language at three different, public locations, where the robot was running continuously for several days. The TAM2 questionnaire results reveal that on average users were motivated to use the robot and found it rather useful and easy to use regardless of the specific location. However, public spaces like the tourist information and the city library seem to be a better fit for our interactive, robotic head than an office environment such as the building authority, where the willingness to interact was lower. Overall, the robot’s multi-lingual responses were very much appreciated, but every fifth user found the response time too slow impeding the dialog flow, which remains to be improved in future work.

[HC-13] owards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging

链接: https://arxiv.org/abs/2607.24081
作者: Parth G. Dangi,Yogesh Kumaar Meena
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)
备注: 6 pages, 3 figures, 2 tables, selected to be presented at Brain-Machine Interface (BMI) Systems Session, IEEE International Conference on Systems, Man, and Cybernetics (IEEE SMC 2026)

点击查看摘要

Abstract:Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challenges in developing BMIs is expanding their usability and control, which can be achieved by accurately decoding multiple kinematic and kinetic parameters. To address this, we propose three regression models: partial least squares regressor, multilayered perceptron, and attention based regressor, to decode multiple movement parameters from EEG signals. We evaluated these models on the WAY EEG GAL dataset, focusing on their performance under subject specific and subject independent conditions with two strategies: a single model for all parameters and a baseline with separate models for each parameter. Among all regressors, the attention based regressor achieved the best performance, with an R^2 of 0.8 and a latency of 29.2 milliseconds, demonstrating significant improvement in simultaneous multi parameter decoding. However, its performance dropped for single parameter decoding. The multi layered perceptron showed more consistent but lower accuracy across both decoding types ( R^2 = 0.49). These findings highlight the potential of attention based models for real time multi command BMI systems and contribute to the development of more intuitive control devices.

[HC-14] he Effect of Photorealism Consistency between the Virtual Hands and Environment on the Sense of Body Ownership and Presence in Virtual Reality

链接: https://arxiv.org/abs/2607.24047
作者: Nami Ogawa,Takuji Narumi
类目: Human-Computer Interaction (cs.HC)
备注: 14 pages, 4 figures. Accepted author manuscript of an article published in IEEE Transactions on Visualization and Computer Graphics. DOI: https://doi.org/10.1109/TVCG.2026.3686249

点击查看摘要

Abstract:Virtual reality (VR) technology allows users to feel a virtual body as if it was their own (i.e., body ownership illusion). Previous studies have explored how the visual realism of the user’s virtual body influences the sense of body ownership in VR, focusing on the aspect of anthropomorphism. However, the effect of photorealism, another element that characterizes the visual realism of a virtual human, has not been systematically examined in the context of body ownership illusion. Therefore, we investigate the effect of the rendering style of virtual hands on the sense of body ownership, hypothesizing that the effect is affected by the rendering style of the virtual environment. In addition, we examine the effect of photorealism on presence (i.e., the sense of being there), as existing studies have offered inconsistent evidence. To this aim, we conducted a 3 x 3 mixed-design remote VR experiment (N=117) that factored in the rendering styles of both the virtual hand and the environment and analyzed the subjective data from the questionnaire. The results suggest that neither the rendering style of virtual hands nor the consistency of the rendering style between the virtual hand and the environment influenced the sense of body ownership. Nevertheless, the more photorealistically the virtual environment was rendered, the stronger the presence. Our results indicate that different dimensions of virtual body realism affect the sense of body ownership differently, thereby highlighting the importance of applying a suitable classification of realism to understand the body ownership illusion for virtual bodies more in depth.

[HC-15] A Design Space for Quantum Circuit Visualizations IEEE-VIS2026

链接: https://arxiv.org/abs/2607.24042
作者: Hyeok Kim,Leilani Battle
类目: Human-Computer Interaction (cs.HC)
备注: Accepted to IEEE VIS 2026 Full Paper, 11 pages, 13 figures

点击查看摘要

Abstract:Quantum circuit visualizations play an essential role in supporting sense-making and communication of quantum programs. While several tools exist for rendering quantum circuits, they vary widely in encoding options due to idiosyncrasies among machine and platform providers. We observe an opportunity to coalesce these disparate rendering approaches under a single, unified grammar to enable consistent, cross-platform enhancement of quantum circuit visualizations. However, it is unclear how to design such a grammar to best support the quantum computing community. Towards this end, we contribute a design space of quantum circuit visualizations by analyzing 182 static and 12 interactive cases collected from online tutorials and documentations, research publications, public presentations, and prior systems. Based on our analysis, we discuss how our design space relates to existing visualization principles yet exhibits unique aspects. We conclude with opportunities for future systems regarding data structure, cognition, and integrability.

[HC-16] A Cyclic Adaptation-Generalization Framework with Uncertainty-Guided Self-Paced Learning for Long-Term Brain-Machine Interfaces

链接: https://arxiv.org/abs/2607.24031
作者: Jiyu Wei,Di Hong,Zhanjie Zhang,Dazhong Rong,Qinming He,Yueming Wang
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Robotics (cs.RO); Signal Processing (eess.SP)
备注:

点击查看摘要

Abstract:Brain-Machine Interfaces (BMIs), which link the brain to external devices, hold great potential in rehabilitation, human performance augmentation, and human-centered robotics. However, invasive BMIs face a critical challenge for long-term deployment due to neural drift, which degrades decoding performance over time and necessitates frequent recalibration. Existing methods designed to mitigate neural drift typically rely on either domain adaptation (DA) or domain generalization (DG) alone and often fail to capture fine-grained distribution shifts across neural subdomains, resulting in limited performance. To overcome these limitations, we propose Uncertainty-guided Self-paced Cycling (UnSPC), a robust framework that synergizes DA and DG for target domain refining under an Uncertainty-guided Self-paced Pseudo-labeling (UnSPL) mechanism. To handle subdomain neural drift across domains, UNSPL is proposed to iteratively mine reliable pseudo-labeled samples with a noise-robust ranking strategy for further fine-tuning. Leveraging these high-quality samples, we introduce a novel Cycling Adaptation and Generalization (CycAG) strategy, which integrates DA and DG within an iterative cycle to progressively mitigate both global and subdomain drift. This cyclic process enables effective alignment to evolving target distributions while preserving robust and transferable representations, thereby mitigating performance degradation under long-term neural drifts. Extensive experiments on multiple neural decoding datasets demonstrate the effectiveness and robustness of UnSPC. To our knowledge, our proposed UnSPC is the first to cyclically integrate DA and DG with pseudo-labeling, paving the way toward stable long-term BMI controls.

[HC-17] Self-Supervised Consistency Enhanced Disentangled Learning for Neural Decoding Generalization in Brain-Machine Interface

链接: https://arxiv.org/abs/2607.24023
作者: Jiyu Wei,Di Hong,Zhanjie Zhang,Dazhong Rong,Qinming He,Yueming Wang
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Robotics (cs.RO); Signal Processing (eess.SP)
备注:

点击查看摘要

Abstract:Brain-Machine Interfaces (BMIs) provide a direct communication pathway between the brain and external devices, enabling humans to control assistive and robotic technologies, with potential applications in rehabilitation, human motor augmentation, and human-centered robotics. However, due to neural drift, the performance of BMIs decreases over time, posing challenges for long-term viability, particularly for invasive BMIs (iBMIs). Existing solutions suffer from two main drawbacks: (i) difficulty in learning robust neural representations, and (ii) neglecting that neural drift varies across motor parameters (e.g., velocity, direction, and speed). To overcome these limitations, we propose Self-Supervised Consistency enhanced Disentangled Learning (SSCDL), a neural decoding generalization framework built on two key innovations. We first design a backbone model named Consistency enhanced Neural Decoder (CND), using a novel teacher-student consistency constraint with simulated neural signal perturbations to learn robust representations invariant to neural drift. Then, we employ three dedicated CNDs under the Complementary-Disentangled Generalization (CDG) mechanism, which disentangles motor signals into velocity, direction, and speed with inspiration from neural preference theory. This disentangled learning enables SSCDL to capture invariant neural representations from diverse neural preference perspectives, significantly enhancing cross-day generalization. Extensive experimental results show that SSCDL delivers state-of-the-art decoding performance, exhibiting high robustness and cross-day stability. These capabilities underscore its strong potential for long-term interaction in human-centric robotic and fine-grained assistive applications.

[HC-18] On Capturing the Narrative: Social Media Manipulation Wargaming for Cyberliteracy

链接: https://arxiv.org/abs/2607.23993
作者: Alexandra Vassar,Rahat Masood,Hammond Pearce
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 7 pages, 1 figure. Paper is under review

点击查看摘要

Abstract:Misinformation is deeply embedded in online discourse, with nearly one in five posts during global events generated by bots that amplify false content. In recent years, the use of Generative AI has further lowered the barrier to producing convincing misinformation, yet most digital literacy education still relies on static checklists and single-player inoculation games built for an earlier media landscape. This paper describes how we addressed this educational gap through Capture the Narrative, a four-week multi-university competition in which student teams build LLM-powered bots to influence a simulated election. We report on our custom social-media platform, the competition environment and design of its 4,000 AI-driven Non-Player Character (NPC) citizens, and what running Capture the Narrative at scale actually involved. In our first iteration, 108 teams from 18 Australian universities produced 7,068,206 player-bot posts, approximately 60% of all platform content. We surveyed 256 students before and 83 after the competition to understand their perceptions of misinformation and the game itself and found that students did not become more confident at spotting bots, contrary to what inoculation theory predicts. Because engagement was rewarded, most teams prioritised high-volume posting over nuanced influence, mirroring real-world platform dynamics. We close with recommendations for educators considering similar interventions, and propose future improvements, such as including a blue-team defensive phase.

[HC-19] Mapping the Reddit Bot Ecosystem: Taxonomy and Evolution

链接: https://arxiv.org/abs/2607.23941
作者: Qiusi Sun,Thomas Gaskin,Branko Blagojevic,Milena Tsvetkova
类目: ocial and Information Networks (cs.SI); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 19 pages, 7 figures, supplementary material included

点击查看摘要

Abstract:Automated agents increasingly participate in online communities, yet their population structure and roles remain poorly understood. Using a dataset of 3,389 identified bots and their full activity histories, we construct a taxonomy of bot “species” on the news aggregation and social media platform Reddit based on temporal, community, linguistic, and semantic features. Clustering analysis reveals 18 distinct bot types spanning content-specialized, behavior-driven, and infrastructural roles such as moderation and utility support. In addition, temporal analysis shows that bot numbers and activity expanded rapidly before peaking around the COVID-19 period, then started declining even before Reddit’s 2023 API policy changes. However, the overall diversity of bot species has remained remarkably stable. These findings suggest that online bot populations form evolving digital ecosystems.

[HC-20] A Sustainable Remote Access Architecture for Digital Inclusion through the Reuse of Discredited TV-BOX Devices

链接: https://arxiv.org/abs/2607.23935
作者: Italo Thiago Felix dos Santos,Carlos Eduardo Correa Queiroz,Adevan Neves Santos,Edgard Luciano Oliveira da Silva
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 12 pages, 3 figures, 3 tables. Published in the Proceedings of the XVIII Brazilian Symposium on Ubiquitous and Pervasive Computing (SBCUP 2026). Original Portuguese title: “Uma Arquitetura de Acesso Remoto Sustentável para Inclusão Digital a Partir do Reaproveitamento de Aparelhos TV-BOX Descaracterizados”

点击查看摘要

Abstract:Digital inclusion and the volume of electronic waste (e-waste) are major challenges for society, with direct impacts on the educational context. Educational institutions, especially those with limited resources, face difficulties in expanding and maintaining computer laboratories. This paper presents the implementation of a sustainable, low-cost Desktop Virtualization Infrastructure (VDI) developed through the reuse of discarded electronic equipment. The proposed solution employs a refurbished Linux server and repurposes confiscated TV Box devices (originally intended for disposal) into functional ARM-based thin clients by installing a compatible Linux distribution. This approach is aligned with the principles of the circular economy by promoting hardware reuse and seeks to reduce the costs associated with deploying computing environments for education. The paper describes the system architecture, the device repurposing process, and the observed results, indicating that this approach represents a viable alternative to support digital education initiatives with a focus on sustainability.

[HC-21] SHARE: Towards Head-Mounted AR with User-Centric SLAM in Shared Human-Robot Workspaces

链接: https://arxiv.org/abs/2607.23901
作者: Tianyuan Du,Tianyi Hu,Hanting Ye,Maria Gorlatova
类目: Human-Computer Interaction (cs.HC); Robotics (cs.RO)
备注: 28 pages, 15 figures, 1 table

点击查看摘要

Abstract:Human-Robot Collaboration (HRC) in shared physical spaces using Augmented Reality (AR) interfaces is powered by Simultaneous Localization and Mapping (SLAM). Existing multi-agent SLAM systems rely on an edge server to combine visual findings of multiple resource-constrained agents, perform computation, and schedule updates to their local maps. However, the edge treats all agents uniformly and ignores the fundamentally different latency requirements of heterogeneous HRC agents: robots and head-mounted AR users. This uniform resource allocation often results in high lag for user manipulation, as it does not meet the stringent latency requirements of AR. In this work, we design, implement, and evaluate SHARE, a user-centric SLAM system that strategically prioritizes AR user experience while maintaining accurate tracking performance for robots. SHARE builds a first-of-its-kind experience model for HRC agents and adaptively adjusts transmission priorities to match it. To reduce end-to-end latency, SHARE leverages the redundancy of visual features acquired by agents in shared human-robot workspaces to reduce computation time induced by edge-based processing. Real-world deployment with commercial AR headsets and a ground robot achieves 13.22 ms average latency for AR users (43.3% reduction from baseline) while maintaining sub-2-centimeter tracking accuracy. User studies further reveal statistically significant improvements in user perception.

[HC-22] Evaluating Closed-Loop EEG Feedback for Simulated Prosthetic Vision in Immersive VR: A Sham-Controlled Feasibility Study

链接: https://arxiv.org/abs/2607.23889
作者: Ruyi Cao,Lily M. Turkstra,Adyah Rastogi,Michael Beyeler
类目: Human-Computer Interaction (cs.HC); Neurons and Cognition (q-bio.NC)
备注:

点击查看摘要

Abstract:Visual prostheses require users to interpret sparse and distorted artificial percepts through active visual search. We developed an EEG-guided neuroadaptive training platform for simulated prosthetic vision in immersive virtual reality and evaluated its feasibility in a sham-controlled object-localization task. Twenty-two sighted participants searched a virtual desk scene rendered through a low-resolution phosphene simulation while EEG was recorded using a dry-electrode headset integrated with a head-mounted display. During training, participants received post-trial visual feedback based either on a commonly used EEG engagement index, \beta/(\alpha+\theta) , or on visually matched non-contingent sham values. Both groups showed comparable within-session improvements in localization performance, consistent with practice, increasing familiarity with the simulated percepts, or refinement of search strategies. EEG-contingent feedback did not produce reliable group-level benefits in localization accuracy, completion time, workload, or modulation of the targeted index. Exploratory analyses showed substantial individual variability, but did not establish a feedback-specific relationship between the engagement index and behavioral performance. These findings demonstrate the feasibility of integrating EEG-contingent feedback with immersive simulated prosthetic vision, while identifying important limitations of the EEG measure, single-session training protocol, and post-trial feedback design.

[HC-23] Visible to the Court: How AI Is (and Isnt) Litigated in U.S. Federal Court Opinions

链接: https://arxiv.org/abs/2607.23888
作者: Julie Yu,Rock Yuren Pang,Jevan Hutson,Katharina Reinecke
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 15 pages, 5 figures

点击查看摘要

Abstract:In the United States, artificial intelligence (AI) is rapidly deployed amid limited federal regulation. With courts become a recurring forum in which AI-related practices are scrutinized, it is important to empirically understand the AI litigation landscape to date. We address this gap through a systematic review of 559 U.S. federal court opinions in which AI plays a role in the parties’ contentions, taxonomizing (1) common topics of dispute, (2) the AI technologies implicated, and (3) the parties involved, including common plaintiff and defendant types. We identify seven recurring dispute areas, six categories of AI technologies at the center of litigation, and four types of common litigants, alongside legal doctrines used by the litigants. A comparison of this taxonomy to the AI Incident Database revealed substantial gaps in coverage, definitions, and prevalence between documented and litigated harms, suggesting courts capture only part of the AI risk landscape. In addition, we found that court decisions primarily rely on pre-existing legal doctrines to manage AI rather than making new AI-specific laws, producing a form of “piecemeal” AI governance. As a result, federal court outcomes are shaped less by where AI has caused harms and more by which harms are cognizable under existing statutes, leading to certain AI harms remaining unresolved.

[HC-24] Meandering Photo Lines: Fluid Multiscale Exploration of Photographic Archives IEEE-VIS2026

链接: https://arxiv.org/abs/2607.23769
作者: Mark-Jan Bludau,Christian Tominski,Marian Dörk
类目: Human-Computer Interaction (cs.HC)
备注: 11 pages, 7 figures. Accepted at IEEE VIS 2026; to appear in IEEE TVCG in 2027

点击查看摘要

Abstract:We introduce a visualization that interweaves spatial and temporal structures of photo collections into explorable pathways. Existing photo browsing interfaces typically use space and time separately for sorting, filtering, or mapping photos. What remains missing is an integrated visual form that represents the photographic journey itself. This is particularly relevant for the individual photographer’s archive, where their unique lifeline is written into the collection, and it is this spatiotemporal trajectory that gives necessary context and meaning to the work. Inspired by the winding nature of rivers, fluid interaction, and un/foldable visualizations, Meandering Photo Lines visualizes adjustable pathways through a photo collection across space, time, and topics. Two complementary views enable visual exploration at different scales: A grid-based arrangement of photo clusters offers a more analytical overview, while meandering spatiotemporal curves facilitate more serendipitous and reflective exploration. Animated transformations connect these two views seamlessly for multiscale navigation from spatiotemporal cluster signatures down to chronological photo series and individual photographs. We demonstrate our approach through two case studies: an archive of approximately 100,000 photographs documenting the Jewish diaspora life worldwide from 1978 to 2021 and a personal smartphone collection of around 20,000 pictures. Qualitative feedback from a think-aloud study indicates that our approach supports both analytical examination and curiosity-driven engagement.

[HC-25] Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

链接: https://arxiv.org/abs/2607.23670
作者: Aayush Kumar,Avik Dutta,Sumit Gulwani,Gustavo Soares,Advait Sarkar,Emerson Murphy-Hill
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execution. However, it remains unclear whether the benefits of this feature translate to end-user programming environments such as spreadsheets. Since spreadsheet programmers tend to work iteratively and care less about technical correctness, upfront planning may not fit into their workflows as easily. In this paper, we build a prototype of a Plan Mode for spreadsheet programming and evaluate it against a non-planning baseline through a within-subjects user study (N=24). We found that despite similar task outcomes with both tools, using Plan Mode led to a reduction in refinement and a better perception of the tool across dimensions of creativity support and human-machine collaboration. We discuss the implications of these results for the future design of Plan Modes, and for the broader role of human-AI planning in end-user programming.

[HC-26] Multimodal Data Comprehension: Understanding How Visual-Textual Chains of Information Influence Data Interpretation

链接: https://arxiv.org/abs/2607.23489
作者: Arran Zeyu Wang,Fuling Sun,Danielle Albers Szafir
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Visualizations and text often work together to support effective data communication. Despite this common paradigm, we know little about how the interplay of these modalities affects people’s data comprehension. We present a novel experimental paradigm to investigate multimodal data comprehension—the process of people comprehending information from multimodal visual and textual data—across both crowdsourced and think-aloud environments. Our methodology employs two sequential chains for presenting multimodal information—a visualization-first chain and a text-first chain—asking people to describe the data presented iteratively. By comparing how people’s data comprehension changes across the chain, we can assess the information contribution of each modality and how they shape subsequent comprehension. We found that the visualization-first chain facilitates exploratory comprehension with hypothesis-driven discovery, whereas the text-first chain yields confirmatory comprehension akin to framing effects where visualizations serve to reinforce and confirm observations drawn from text. Our findings provide empirical insights into multimodal information integration, with implications for designing more effective data-driven communication.

[HC-27] Visualizing in the Minds Eye: Icon Design Shapes Mental Imagery of Fire Risks IEEE-VIS2026

链接: https://arxiv.org/abs/2607.23369
作者: Wen Xu,Anjana Arunkumar,Lace M. Padilla
类目: Human-Computer Interaction (cs.HC)
备注: 5 pages, 3 figures, accepted to IEEE VIS 2026 (short paper)

点击查看摘要

Abstract:We introduce mental imagery, or seeing images in the “mind’s eye,” as a cognitive process that can be shaped by data visualization design and in turn impact decisions. We found in a preregistered study (n = 400) that abstract geometric icons visualizing fire risk data evoked more mental images than concrete house-on-fire icons, and also produced more diverse and personalized mental images. Mediation analysis showed that increased mental imagery subsequently led to risk-averse decisions through evoking negative affect. These findings reveal a nuanced mechanism through which visualization concreteness influences decisions: concrete designs may actually suppress affect-driven behavior by restricting mental imagery.

[HC-28] Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models

链接: https://arxiv.org/abs/2607.23309
作者: Shuoyang Jasper Zheng,Anna Xambó Sedó,Nick Bryan-Kinns
类目: Human-Computer Interaction (cs.HC); Sound (cs.SD)
备注: Under review for “Explainable AI for the Arts” (N. Bryan-Kinns, Ed.), Springer

点击查看摘要

Abstract:Recent work in Human-Computer Interaction (HCI) increasingly treats AI models as design materials that have distinctive computational properties to shape design artifacts. Artists learn to work with the model “at play” to explore their emerging properties. The aim of explainability, in this view, is to make visible a crafting and hacking space to enable sustained creative practices with AI. In this chapter, we propose material explainability as a range of activities and artifacts that transform AI models into accessible and inclusive design materials in the workspace of artists, designers, and makers. We present a case study of building a repository of resources to enable artistic explorations of neural audio models in New Interfaces for Musical Expression (NIME) design. Reflecting on our community-building journey and the making of a collection of musical interface designs with a group of artists, we raise three recommendations on enabling the exploration of AI as materials in artistic practices to inspire future XAI design for artists.

[HC-29] Cross Sensory Co-Design Tools and Interaction Qualities

链接: https://arxiv.org/abs/2607.23298
作者: Albrecht Kurze,Klaus Stephan
类目: Human-Computer Interaction (cs.HC)
备注: In “Workshop Cross-Sensory Futures: Rewiring Perception in HCI” at CHI 2026. April 13, 2026. Barcelona, Spain

点击查看摘要

Abstract:What is the temperature of a loud cry? What is the color of a hand-clap? Cross-sensory and multi-modal interactions offer interesting possibilities in terms of inputs and outputs that might be leveraged for affective and emotional purposes. However, it is hard to design such interactions without the right co-design tools. We developed such tools in the last 10 years: the Loaded Dice and the Wheel of Plush. Both devices incorporate a multitude of sensors and actuators to demonstrate interactions and aid in co-design workshops. The tools incorporate the principle of technical synesthesia: the ability to map any of the included sensors to any of the included actuators. With intuitive simple interaction schemes it is possible to easily and effortlessly create new combinations. We discuss the core principles of the tools and how we realized a meaningful mapping between sensors and actuators. We further discuss how we use the tools for exploration, sensory sketching or scenario driven ideation, and where we see additional potential.

[HC-30] A Taxonomy of Confabulations and the Perception-Reality Gap in LLM -Assisted Immersive Scene Editing

链接: https://arxiv.org/abs/2607.23213
作者: Junlong Chen,Per Ola Kristensson
类目: Human-Computer Interaction (cs.HC)
备注: Under review

点击查看摘要

Abstract:Large language models (LLMs) are being increasingly integrated into immersive environments and design workflows, providing application prospects in areas such as rapid scene prototyping for non-expert users and scene understanding capabilities for accessibility design. While many workflows that incorporate LLMs in immersive spaces are proposed, such systems can exhibit errors, potentially resulting in frustration, loss of user trust, and compromised user safety. This paper studies the underexplored area of LLM confabulations in immersive 3D scene editing contexts. Through an exploratory study with 24 non-expert users, we construct a taxonomy of the different types of confabulation observed in LLM-assisted immersive 3D scene editing. We report their prevalence and disruptiveness, and define the construct perception-reality gap to help understand the gap between the actual and perceived occurrence of confabulations. We highlight the observed saturation of confabulation awareness under load and conclude by discussing design implications for confabulation mitigation in future LLM-assisted systems in immersive 3D scenes.

[HC-31] Beyond Conversations: Spatially-Anchored Previews for Intent Disambiguation in LLM -Assisted Geometry Editing in Virtual Reality

链接: https://arxiv.org/abs/2607.23201
作者: Junlong Chen,Amr Gomaa,Jens Grubert,Per Ola Kristensson
类目: Human-Computer Interaction (cs.HC)
备注: Under review

点击查看摘要

Abstract:User intent disambiguation remains a key challenge in intelligent interactive systems. While they have been widely studied in dialogue systems in 2D interfaces, research on how intent disambiguation could be incorporated within Large Language Model (LLM) assisted editing workflows in immersive environments remains limited. Recent advances in LLMs create opportunities to leverage the immersive nature of virtual and augmented reality (VR/AR) environments to provide better disambiguation support. In this paper, we evaluate how traditional dialogue-based disambiguation can be augmented with spatially-anchored graphical previews to resolve ambiguous user commands in LLM-assisted parameter-driven editing workflows. A within-subjects study in which 24 participants completed complex geometry editing tasks in VR simulate scenarios where VR scenes are controlled by numerical parameters. Compared with the condition where disambiguation is not available, quantitative metrics and qualitative feedback indicate that a hybrid approach which combines clarification questions and graphical previews can support better interaction stability with fewer conversation rounds while improving user experience. These findings provide empirical evidence on the effectiveness of disambiguation methods in LLM-assisted editing of parameter-driven immersive scenes and inform design guidelines for future integration of LLMs in advanced VR/AR systems.

[HC-32] From Vibe to Code – and Back: Lexical Oscillation in the Formation of Design Intent with Generative AI

链接: https://arxiv.org/abs/2607.23126
作者: Daisaku Sato
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: Preprint. Submitted to the International Journal of Human-Computer Interaction (IJHCI) and currently under review

点击查看摘要

Abstract:Generative AI design tools make natural-language prompts a starting point for design, placing new articulation demands on designers. Rather than treating prompts as the transmission of pre-existing design intent, we ask how design intent is formed through situated interaction with AI. Five expert UI/UX designers (11-20 years’ experience, M = 15.4) designed landing-page hero sections with a generative AI tool, recorded through think-aloud and retrospective interviews. Using reflexive thematic analysis, we used lexical granularity (L1 vibe, L2 design-domain, L3 operational language) as a sensitizing lens. Rather than moving from vibe to code unidirectionally, designers showed lexical oscillation, including returns from operational specificity to ambiguity. Mismatches with AI outputs were taken up as occasions for designers to reconsider what they meant, and engagement shifted from instruction to consultation. One non-oscillating trajectory – a negative case – suggested conceptual misalignment as a tentative boundary for future examination. We position AI as a non-neutral generative interlocutor and ambiguity as a resource for design judgment.

[HC-33] ouching or Chatting: The Utility of LLM s and Tactile Charts for Learning about Complex Chart Types by BLV Individuals

链接: https://arxiv.org/abs/2607.23065
作者: Tingying He,Maggie McCracken,Daniel Hajas,Sarah Creem-Regehr,Alexander Lex
类目: Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Visualizations are central to communicating data, yet blind and low-vision (BLV) people often lack support for understanding chart types—knowledge that is essential for interpreting new visualizations and collaborating with sighted peers. Prior work found that BLV individuals viewed example tactile charts as more helpful than text-only approaches and preferred them for learning advanced chart types, particularly for understanding spatial layouts and shapes. Meanwhile, large language models (LLMs) are increasingly used by BLV individuals for chart explanation and question answering (QA), but have been studied primarily for dataset exploration rather than chart-type learning. Existing LLM-based chart QA also shows that users frequently ask about layout and structure, yet struggle with spatial concepts and misdirect questions when mental models are weak. We investigate how LLMs influence chart-type learning and whether tactile learning improves subsequent LLM-supported exploration. We extend our tactile chart learning tools with an LLM chatbot that provides interactive explanations and supports follow-up questions. In an interview study with 12 BLV participants, we compare two learning formats: (1) a tactile chart, a textual explanation, and an LLM chatbot; and (2) a textual explanation and an LLM chatbot. The learning phase was followed by exploration of an unfamiliar dataset using alt text and an LLM. Thematic analysis shows that tactile templates support BLV participants’ formation of chart-type mental models, which scaffolds subsequent LLM-mediated data exploration. Text+LLM explanations without tactile support show weaknesses for spatial-reasoning tasks.

[HC-34] he Help Ladder: Skill-Adaptive Peer Scaffolding for Real-Time Collaborative Programming

链接: https://arxiv.org/abs/2607.23031
作者: Panayu Keelawat,Sriram Narlapati,Sangwook Lee,Darshan Nere,Griffin Ogura,Xinran Adeline Li,Sang Won Lee,Yan Chen
类目: Human-Computer Interaction (cs.HC)
备注: Accepted to VL/HCC 2026

点击查看摘要

Abstract:Collaborative programming is a widely adopted classroom activity to encourage peer scaffolding, yet real-time collaboration often breaks down into parallel individual work with minimal interaction. Our formative studies reveal that even when students want to collaborate, they are held back by the effort required to understand a teammate’s entire problem at once. We present Canary, a system that supports peer scaffolding by breaking down programming obstacles into smaller steps tailored to a student’s skill level. Canary alerts potential helpers to specific places where they can start, using AI to turn complex problems into a step-by-step ladder that starts with easy fixes before moving toward harder logic. By providing this gradual ramp-up, Canary enables students to make quick contributions and progressively work toward solving their peers’ problems. Our evaluation shows that this staged approach makes helping feel less overwhelming, leading to more frequent and effective collaboration among students.

[HC-35] WCM: World-Cognition Model for Generalizable Human-Robot Interaction

链接: https://arxiv.org/abs/2607.22999
作者: Yuzhen Chen,KC Zhou
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tasks. Current robot-control paradigms, including vision-language-action policies and world-model-based planners, are mainly optimized for instruction execution, leaving users with little visibility into why an action is chosen and few mechanisms to redirect, correct, or teach the robot through interaction. To solve this problem, we present the World-Cognition Model (WCM), a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime. SLAK separates perception, reasoning, control, and memory, while the runtime allows reasoning, dialogue, and execution to proceed concurrently. WCM further introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks. Teaching episodes and autonomous task rollouts are refined into chain-of-thought supervision to continually improve the model. WCM achieves a 73.8% average success rate across nine real-world human-robot interaction tasks, including tasks held out from CoT fine-tuning and a long-horizon task learned through teaching.

[HC-36] Pose-Aware Modeling to Mitigate Pose-Related Artifacts in Tactile Gloves

链接: https://arxiv.org/abs/2607.22964
作者: Tianhong Catherine Yu,Ziyi Kou,Mia Huang,Taylor Niehues,Yiyue Luo,Li Guan,Dingtian Zhang
类目: Robotics (cs.RO); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Tactile gloves digitize contact and force during hand-object interactions, enabling robotics applications in dexterous manipulation, teleoperation, and learning from demonstration. To preserve hand dexterity and capture the nuances of natural interactions, these gloves and the integrated tactile sensors are designed to be soft, flexible, and comfortable. However, such flexible sensors are sensitive not only to contact forces but also unavoidably to hand pose changes, resulting in pose-related artifacts (PRAs). PRAs are especially problematic in the low-force range, resulting in misdetections or late-onset detections of contact, which raises the minimum detectable force (MDF) of the glove. In this work, we characterize the PRAs in relation to pose and force. Building on these insights, we introduce a glove-agnostic algorithmic framework that leverages hand pose information, which is increasingly available, to mitigate PRAs without glove modifications. Our pose-aware force estimation model augments tactile-to-force pipelines with a residual prediction branch that explicitly accounts for pose-induced sensor deformations. We validate our approach across 3 glove designs and 15 users, reducing MDF by 10.4%, 12.2%, and 18.3%, with consistent improvements across all evaluated metrics. This method provides a practical path to improving the usability of tactile gloves in data collection and diverse robotic applications.

[HC-37] “Why SuaCode?”: Understanding African Students Motivations for Taking a Smartphone-Based Online Coding Course

链接: https://arxiv.org/abs/2607.22940
作者: Michael Addo,Nana Maryam Munagah,Victor Kumbol,Judith Uchidiuno,George Boateng
类目: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY)
备注: Accepted as poster at ACM COMPASS 2023

点击查看摘要

Abstract:Computer programming MOOCs are instrumental in providing students with high-quality instruction in areas where there is limited access. They are especially beneficial to post-secondary African students as less than 1% of them leave secondary school with fundamental coding skills. One strategy for increasing their efficacy for African students is to understand students’ motivation for enrolling. These insights can inform the design of MOOC content and assessments to align with students’ interests. We administered an open-ended response survey to (self-identified) Africans enrolled in a smartphone-based online coding course (SuaCode). We analyzed a random sample of 450 (of 3000) responses using a grounded theory approach. We found that most African students (68.7%) participated in SuaCode for intrinsic reasons such as improving themselves, learning with like-minded individuals, and gaining skills to help address societal issues. We discuss the implications of these findings in the design of programming MOOCs targeted at African students.

[HC-38] Reflections and Recommendations on AI Adoption Practice from a Mixed-Ability Research Group

链接: https://arxiv.org/abs/2607.22886
作者: Shalini Madan,Sreelakshmi Surabiyil Bindu,Veronica Pimenova,Ellie Seehorn,Venkatesh Potluri
类目: Human-Computer Interaction (cs.HC)
备注: Submitted and accepted atThe 28th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS '26), October 25–28, 2026, Vila Nova de Gaia, Portugal

点击查看摘要

Abstract:Generative AI tools have recently been rapidly adopted by academics in mixed-ability research teams for both personal and professional tasks. While previous work on adoption of AI-based workflows has focused on collaboration and productivity, the perceptions of AI use within research teams remains divided. Through qualitative analysis of interviews of the five members of our mixed-ability research team, we discuss the motivations, challenges, and practices surrounding the use of generative AI in our lab. We reflect on experiences that shaped recommendations for balanced AI use that enable mixed-ability team workflows: (1) managing disability tax crip time, (2) homogenizing identity, (3) risk disclosure of private information, (4) self-experimentation and miscellaneous tasks, and (5) information seeking. We build upon these themes to present AI practice recommendations we established for our lab to promote AI workflow adoption while preserving agency and disability identity.

[HC-39] Physical AI Governance: From Theory to Practice Across Life Cycle

链接: https://arxiv.org/abs/2607.22877
作者: Wang Yang,Shaobo Wang,Hongxuan Liu,Xiaoran Cai,Yunyu He,Jingzong Zhou,Mengzhong Ma,Yi Yu,Rohit Sharma,Jingjing Fu,Peng Qi
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact with, and act in the physical world. Unlike traditional AI, Physical AI operates under real-time safety constraints, continuously interacts with dynamic environments, and coexists with humans, introducing governance challenges that existing AI governance frameworks do not explicitly address. This paper presents a comprehensive survey of Physical AI governance from both scientific and operational perspectives. We synthesize existing governance principles and organize them into a unified governance framework tailored to physical AI systems. Building on this foundation, we propose a five-stage Physical AI lifecycle comprising research, design, data, model development, and deployment, and demonstrate how governance can be operationalized across each stage through concrete implementation practices. By connecting governance principles with engineering workflows, this survey provides a structured reference for researchers, developers, and policymakers to build Physical AI systems that are safe, trustworthy, and aligned with societal values.

[HC-40] Validation of a Real-Time Manual Wheelchair Simulator through Biomechanical and Perceptual Measures: A Comparison with Overground Propulsion

链接: https://arxiv.org/abs/2607.22851
作者: Ateayeh Bayat,Félix Chénier
类目: Human-Computer Interaction (cs.HC); Biological Physics (physics.bio-ph)
备注: 18 pages, 2 figures

点击查看摘要

Abstract:Purpose: Manual wheelchair simulators can provide a safe and controlled environment for propulsion training, biofeedback, and biomechanical assessment. However, their ecological validity must be established to ensure that simulator-based outcomes reflect overground propulsion. This study aimed to evaluate the ecological validity of a real-time manual wheelchair simulator by comparing biomechanical and perceptual outcomes during matched overground and simulator tasks. Materials and Methods: Thirteen participants (10 able-bodied individuals, 3 experienced manual wheelchair users), performed propulsion tasks overground and on the simulator. Task segments included straight-line propulsion, acceleration, turning, ascending, and descending. Biomechanical outcomes were collected using instrumented wheels and compared between environments using temporal, kinetic, and waveform-based measures. Perceived realism and user experience were assessed using task-specific realism ratings, a post-test questionnaire, and semi-structured interviews. Results: Temporal variables showed relatively small differences between overground and simulator propulsion, whereas kinetic variables showed larger discrepancies, consistent with previous simulator comparisons. Waveform patterns for force, moment, and power demonstrated high similarity between environments. Participants generally reported positive perceptions of realism, safety, and satisfaction, and highlighted the value of practicing wheelchair skills in a controlled environment. Conclusion: The simulator demonstrated partial ecological validity, particularly for temporal and waveform-based propulsion outcomes. Its ability to accommodate the user’s own wheelchair configuration may help preserve ergonomics and enhance realism, supporting its potential use as an assistive technology tool for manual wheelchair propulsion training.

[HC-41] GeoTEAM: A Geospatial Tangible User Interface for Exploration and Visual Analysis of Migration Data

链接: https://arxiv.org/abs/2607.22825
作者: Karen Penaranda Valdivia,Nujaimah Ahmed,Aswah Butt,Nao Nagashima,Sarah Hoyos-Hoyos,Krupa Shah,Elise van den Hoven,Robert McLeman,Roozbeh Manshaei,Gabby Resch,Jamy Li,Ali Mazalek
类目: Human-Computer Interaction (cs.HC)
备注: Submitted to IEEE for possible publication

点击查看摘要

Abstract:Migration studies is an interdisciplinary field, requiring collaboration between researchers with varying levels of geospatial and quantitative data literacy. Geospatial tangible user interfaces (GTUIs) offer promising opportunities for embodied and collaborative spatial exploration of migration data. Few GTUIs, however, provide real-time visual feedback of data values to help users collaboratively identify trends. To address this gap, we present GeoTEAM, a novel tangible system for real-time, dynamic regional exploration, co-designed with migration and HCI researchers. GeoTEAM features active tangible dials with built-in touchscreens designed to control time and navigate map layers for net migration and various environmental sub-drivers in a multi-surface environment with tabletop and wall displays. We evaluated this system with nine pairs of researchers possessing multi-expertise in geospatial data literacy. Qualitative findings show that our system promotes collaboration, intuitive sense-making, and increased user confidence, enabled by physical interaction and real-time feedback from the active tangibles.

[HC-42] xt-based Tactile Graphics Generation for the Visually Impaired ECCV2026

链接: https://arxiv.org/abs/2607.22674
作者: Ruihan Gao,Joonghyuk Shin,Ava Pun,Jaesik Park,Wenzhen Yuan,Jun-Yan Zhu
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
备注: ECCV 2026, Project webpage: this https URL

点击查看摘要

Abstract:Tactile graphics are a primary medium for blind and low-vision (BLV) individuals to access non-textual information. However, they are difficult to scale or personalize. While recent generative models have revolutionized visual content creation, they are optimized for screen-based visual realism and fail to satisfy the haptic perceptual and physical fabrication constraints required for touch. We present the first integrated generative system that produces fabrication-ready 2.5D tactile graphics directly from natural language prompts, jointly generating global base geometry, fine-grained tactile surface textures, and standard-compliant braille within a unified 3D-printable representation. Our approach introduces fabrication-aware techniques, including template-guided relief generation, a fast diffusion-based text-to-texture module for high-resolution tileable normal maps, and strict base flattening to ensure tactile readability and printability, while supporting both automatic generation and interactive texture control. Extensive evaluations, together with in-person user studies with BLV participants and blindfolded sighted participants using physically 3D-printed outputs, show that participants consistently prefer our results over baselines. By extending generative graphics beyond screens to touchable reliefs, our work broadens access to generative AI for the BLV community and beyond.

[HC-43] A didactical-driven teacher assistant for a dimensional modeling course

链接: https://arxiv.org/abs/2607.22598
作者: Laurent Brisson(IMT Atlantique - DSD),Maria Segarra(IMT Atlantique - INFO, Lab-STICC_MOTEL),Grégory Smits(IMT Atlantique - INFO, Lab-STICC_MOTEL)
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:

点击查看摘要

Abstract:Educational chatbots powered by large language models (LLMs) show promising effects on learning outcomes, yet most systems delegate pedagogical decisions such as content selection and didactic structuring implicitly to the LLM, making tutoring strategies difficult to trace, evaluate, and reproduce. This paper presents a didactical-driven teacher assistant for a French-language university course on dimensional modelling, operating without commercial LLM budget or GPU infrastructure. The architecture formalises the instructor’s pedagogical reasoning into deterministic modules that handle intent detection, concept linking, and didactic approach selection before any text is generated; the LLM acts solely as a linguistic executor. Evaluation on 195 authentic student questions addresses two research questions. First, we show that standard semantic retrieval alone does not reliably recover the pedagogically required content, thereby justifying the upstream orchestration strategy adopted in our architecture (RQ1). Second, compared to free-tier LLMs whose detection performance varies widely across models and which produce errors silently, the deterministic pipeline achieves high pair precision (73%) with full traceability and explicit abstention, though its limited coverage confirms that the detection strategy requires further refinement (RQ2).

计算机视觉

[CV-0] Data Pyramid for Embodied Manipulation

链接: https://arxiv.org/abs/2607.24744
作者: Yifan Ye,Yankai Fu,Yaoxu Lv,Bohan Hou,Jun Cen,Lingdong Kong,Duo Zheng,Tianxing Chen,Jiaming Liu,Ziang Cao,Yunfan Lou,Wei Chow,Xian Sun,Yingshuo Wang,Kuangzhi Ge,Xiaowei Chi,Xidong Zhang,Zhibo Pang,Yiwu Zhong,Sirui Han,Zhihe Lu,Weihao Yuan,Qifeng Chen,Michael Yu Wang,Yao Mu,Ziwei Liu,Jianfei Yang,Ping Luo,Shanghang Zhang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Awesome Embodied Data Pyramid; Project Page at this https URL GitHub Repo at this https URL

点击查看摘要

Abstract:Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a “pyramid” spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.

[CV-1] Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

链接: https://arxiv.org/abs/2607.24731
作者: Bingnan Li,Haozhe Wang,Haozhong Xiong,Fangtai Wu,Jinpeng Yu,Yang Shi,Jiaming Liu,Ruihua Huang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negative-branch errors can compensate in the guided prediction. Through two contrasting cases, we find that naive matching remains effective under shared negative conditioning, where both branch errors decrease jointly. When the model’s native CFG schema retains privileged information in the teacher’s negative branch that is unavailable to the student, however, this joint reduction breaks down and the composed objective induces antagonistic branch-error dynamics, reducing the positive-branch error while increasing the negative-branch error. We term this failure mode Negative Branch Asymmetry (NBA). To address NBA, we introduce Positive–Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction. We apply PDM to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.

[CV-2] KANEx: Translating Kolmogorov-Arnold Networks Interpretability to Medical Explainability MICCAI2026

链接: https://arxiv.org/abs/2607.24730
作者: Krithi Shailya,Ananya Lakshmi Ravi,Venkatanathan K. V.,Sowmya S. Sundaram,Gokul S. Krishnan,Aditi Anand,Balaraman Ravindran
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: MICCAI 2026

点击查看摘要

Abstract:Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In clinical workflows, chest X-ray classifiers are increasingly paired with Vision-Language Models (VLMs) to generate natural-language explanations. However, these systems add linguistic fluency without addressing the underlying opacity of the visual model. With the emergence of Kolmogorov-Arnold Networks (KANs), whose spline-based components provide inherently interpretable functional units, we investigate whether this architectural transparency can be leveraged to produce more trustworthy textual explanations. We introduce KANEx, the first ever framework that leverages the symbolic transparency of KANs to ground VLM reasoning. This interpretability also made it possible to design KAN-Map, a novel heatmap generation method derived directly from KAN models rather than gradient approximations. We feed these grounded contexts into downstream VLMs for enhanced explainability. Benchmarked on the MIMIC-CXR dataset, we demonstrate that KAN-based architectures with ResNet/ViT baselines demonstrate improved semantic similarity while producing significantly more faithful saliency maps. KAN architectures improve visual localization and downstream reasoning quality by 10%. Our findings suggest that grounding linguistic explanations and visual attributions in mathematically interpretable units is a necessary step toward trustworthy medical AI.

[CV-3] MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale MICRO

链接: https://arxiv.org/abs/2607.24729
作者: Huy Huynh,Jingwei Ma,Brian Curless,Ira Kemelmacher-Shlizerman,Steven M. Seitz
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:We introduce MicroZoom, a generative framework for gigapixel image synthesis at the microscopic scale. Given a standard photograph and a sparse set of consumer-grade microscope close-ups, MicroZoom synthesizes a seamless, gigapixel-resolution image grounded in the material character of the real references, enabling exploratory visualization of microscopic texture across the full spatial extent of an object. Our goal is plausible synthesis, not exact reconstruction. We focus on full-image, reference-based, extreme-scale super-resolution at magnification levels of up to 350x, a setting that introduces two major challenges: (1) recovering texture-specific detail from highly lossy inputs near ambiguous material boundaries, and (2) preserving correct large-scale pattern structure, such as the repeating geometry of a fabric weave, across millions of local predictions. We address these with a two-stage cascaded design, where the first stage recovers global pattern coherence and the second refines local texture detail, supplemented by a segmentation mask to guide synthesis at ambiguous boundaries. We verify our approach on a collection of self-captured everyday objects and demonstrate globally coherent, materially grounded gigapixel imagery.

[CV-4] Infrared Imaging Empowered by Artificial Intelligence for Pediatric Skeletal Triage: A Narrative Review and Future Perspectives

链接: https://arxiv.org/abs/2607.24727
作者: Sajad Amiri,Pardis Afshar,Elham Anjomshoa
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 20 pages, 4 figures

点击查看摘要

Abstract:Background. Pediatric musculoskeletal trauma represents up to 18% of pediatric ED visits, yet diagnosis still depends on ionizing radiography. Cumulative low-dose radiation in early life raises lifetime leukemia and brain malignancy risk, motivating radiation-free triage alternatives. Objective. To synthesize evidence for a hybrid framework coupling broad-spectrum infrared (IR) imaging with deep-learning cross-modal translation to generate clinically interpretable synthetic-radiograph reconstructions from non-ionizing data. Approach. We review five IR spectral windows spanning 650 nm to 1 mm - NIR-I, NIR-II, SWIR, MIR/LWIR, and THz - and how dual-geometry (transmission/reflection) acquisition exploits wavelength-specific tissue depth and biochemical sensitivity. We summarize image-to-image translation networks (Pix2Pix, CycleGAN, Swin-Unet) and feature-matching algorithms (SuperPoint, SuperGlue, ALIKED, LightGlue) used to align and fuse IR data into radiograph-equivalent reconstructions. Implications. Pediatric anatomy - smaller cross-sections, thinner cortical bone - favors IR penetration, enabling compact, portable, non-ionizing triage hardware. Feasibility is grounded in fNIRS and transcranial photobiomodulation evidence: near-infrared light passes through skin, skull, and cortex with sufficient signal for hemodynamic monitoring - a longer, more attenuating path than through a pediatric forearm or distal leg. Key barriers: paired IR/X-ray dataset construction, AI-as-medical-device regulatory pathways, generalization across body habitus and skin pigmentation, and acquisition-protocol standardization. Conclusions. Integrated multi-spectral IR+AI imaging is a promising radiation-free complement to pediatric skeletal radiography. Progress requires multi-center paired datasets, externally validated models, and IR source safety qualification under IEC 60825-1. Comments: 20 pages, 4 figures Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.24727 [cs.CV] (or arXiv:2607.24727v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2607.24727 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Pardis Afshar [view email] [v1] Mon, 27 Jul 2026 17:56:53 UTC (5,715 KB)

[CV-5] DreamStyle3D: Efficient 3D Stylized Asset Generation via Dual-Attention Disentanglement ACM-MM2026

链接: https://arxiv.org/abs/2607.24721
作者: Kai Wang,Ziheng Ouyang,Xuying Zhang,Ming-Ming Cheng,Qibin Hou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: ACM MM 2026

点击查看摘要

Abstract:With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. However, existing approaches still struggle to jointly preserve style fidelity, geometric consistency, and generation efficiency, as most of them still rely on indirect 2D-to-3D stylization pipelines. This motivates a native 3D stylization framework that can explicitly disentangle style from geometry while remaining efficient. To this end, we propose DreamStyle3D, an efficient framework for stylized 3D asset generation built on a Decoupled Dual Cross-Attention mechanism. Our method explicitly separates geometric and stylistic features to enable efficient style injection while preserving structural consistency, and further adopts a lightweight training strategy to enhance style consistency and model generalization. In addition, we build an automated data pipeline and construct a dataset of about 15K content-style-stylized triplets for training and evaluation. Extensive experiments demonstrate that our DreamStyle3D can generate high-fidelity, geometrically consistent stylized 3D assets within 10 seconds, substantially improving efficiency while maintaining superior style quality and offering a new solution for 3D content creation. The code and data are available at this https URL.

[CV-6] ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

链接: https://arxiv.org/abs/2607.24707
作者: Ali Ansari,Yasmin Mohammadi,Farnoush Nili,Parsa Esmaeilkhani,Longin Jan Latecki,Eduard Dragut
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Databases (cs.DB)
备注:

点击查看摘要

Abstract:Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUnderstand, the first large-scale benchmark for structured understanding of ER diagrams, comprising 2,960 diagrams collected from curated educational sources, real-world schemas, and synthetically generated examples spanning diverse domains, notations, complexity levels, and Extended Entity-Relationship (EER) constructs. Each diagram is paired with a standardized machine-readable representation for fine-grained evaluation of schema elements. Evaluating state-of-the-art Vision-Language Models (VLMs), we find that while common ERD elements are recovered reliably (F1 0.74), performance drops sharply on weak entities (as low as 0.28 F1), multivalued attributes (0.14 F1), and N-ary relationships (0.07 F1). Reasoning-augmented models improve overall performance by 15-25% but remain sensitive to linguistic priors and increasing diagram complexity. ERUnderstand provides a standardized benchmark for evaluating multimodal understanding of conceptual database schemas. The benchmark, dataset, evaluation toolkit, and generation code are publicly available at this https URL.

[CV-7] SADe: Sparse-Atom Support Decontamination for Few-Shot Segmentation with Weak Support Annotations

链接: https://arxiv.org/abs/2607.24706
作者: Hang Xing,Guangjun Liu,Yan Xia,Xueming Ding
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 8 figures; includes a technical appendix

点击查看摘要

Abstract:Few-shot segmentation (FSS) commonly assumes clean pixel-level support masks, yet practical support supervision often uses boxes, scribbles, coarse masks, or pseudo-masks. These weak annotations may include texture-similar distractors and background context alongside the target, contaminating class prototypes or visual prompts before query prediction. We introduce SADe, a predictor-agnostic support decontamination layer that estimates the reliability of selected support patches without query information. Central to SADe is sparse autoencoder (SAE) atom evidence: dense similarity may respond to both target and texture-similar context, whereas contrasting atom activations inside and outside the weak-support region provides factor-level reliability cues. A lightweight router combines atom evidence with dense similarity and episode statistics to predict patch reliability and generate a cleaned support mask. Trained once on synthetic weak-support episodes from FSS-1000, the router is frozen for all target evaluations. The resulting mask supports standalone prediction or can be supplied to heterogeneous FSS models through native support interfaces without altering query-side inference. Under a matched weak-support protocol, SADe achieves the highest query mIoU in six of nine standalone prompt-shot combinations. With the same ProMi query head, it is within 0.03 mIoU of SAM3-derived masks under tight boxes and surpasses them by 11.17 and 19.49 points under box-r2 and box-r4, respectively. As a plug-in, SADe improves over raw support in 70 of 72 matched box-family comparisons across four frozen downstream models and two datasets. On point and scribble prompts, its average performance remains close to the corresponding raw-support baseline. Ablations and atom-removal controls show that atom evidence contributes reliability information beyond dense similarity.

[CV-8] Panda: Unsupervised Pelvic Anomaly Detection for Real-Time MR Imaging

链接: https://arxiv.org/abs/2607.24703
作者: Anika Knupfer,Maximilian Lindholz,Johanna Paula Müller,Jordina Aviles Verdera,Smiti Tripathy,Susanne Schulz-Heise,Jana Hutter
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 5 figures

点击查看摘要

Abstract:Female pelvic diseases remain an under researched area characterized by often delayed diagnosis. While pelvic MRI offers superior soft-tissue contrast for diagnosis and image-guided procedures, real-time anomaly detection remains challenging due to physiological motion, tissue deformation, and instrument artifacts. Existing supervised approaches are impractical, as adverse events are rare, heterogeneous, and difficult to annotate. We present a Dinomaly-based unsupervised anomaly detection framework adapted for pelvic MRI that learns normative representations from healthy cases and flags deviations without requiring labels. Our approach leverages a frozen DINOv3 Vision Transformer encoder combined with a noisy MLP bottleneck and Linear Attention decoder to prevent identity mapping while maintaining computational efficiency. Anomalies are localized via per-token cosine distance between encoder and decoder representations, yielding spatial anomaly maps that provide immediate feedback at the scanner to support radiologist decision-making and adaptive protocol adjustment. Evaluated on a curated subset of the Uterine Myoma Dataset, the framework achieves a pixel-level AUROC of 88.06% and high specificity (95.45%) at frame level at 40.5 slices/s, meeting real-time clinical deployment requirements. The spatial anomaly maps and frame-level scores provide immediate, localized feedback at the scanner to support radiologist decision-making and adaptive protocol adjustment during active procedures.

[CV-9] Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking CVPR2026

链接: https://arxiv.org/abs/2607.24701
作者: Andong Lu,Ziyi Zha,Jiandong Jin,Shihao Li,Chenglong Li,Jin Tang,Bin Luo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by CVPR2026

点击查看摘要

Abstract:Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be unsatisfactory. In addition, current approaches exhibit limited flexibility in processing both missing and complete data. To overcome these limitations, we propose a Spatio-temporal Conditional Denoising Transformer (SCDT), which integrates the spatial cues and the temporal context to adaptively perform information reconstruction of missing modalities and feature enhancement of weak modalities in a unified framework, for robust modality-missing RGBT tracking. In particular, SCDT leverages the short-term temporal cues from recent historical frames to capture the fine-grained temporal correlations and the long-term temporal cues encoding modality evolution to capture the global context. By jointly exploiting long short-term temporal contexts as the conditions, SCDT progressively guides noisy features of available modalities to learn reliable and temporally consistent multimodal representations. Furthermore, SCDT introduces a noisemodulated adaptation mechanism that dynamically adjusts its behavior according to the modal availability, enabling a single framework to unify feature learning under both modality-missing and complete scenarios without changing the architecture or parameters. Extensive experiments on three public benchmark datasets demonstrate that our method consistently outperforms state-of-the-art methods. The code is available here.

[CV-10] Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification

链接: https://arxiv.org/abs/2607.24683
作者: Francisco Mena,Dino Ienco,Roberto Interdonato,Cassio F. Dantas,Simon Besnard
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at Discovery Science 2026

点击查看摘要

Abstract:Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scenarios subject to operational constraints, such as sensor failures or privacy restrictions, lead to inconsistent modality availability between training and inference times. To handle missing modalities, prior studies have mainly covered bimodal data setups and focused on designing robust fusion processes. Instead, we adopt a multi-modal co-learning framework that prioritizes inter-modal collaboration rather than multi-modal fusion. Specifically, we consider that any subset of modalities may be absent, without assuming predefined missing-modality patterns, an inference scenario we refer to as missing arbitrary modalities. To address this challenge, we introduce two alternative approaches that leverage information at both feature- and decision-level. Experiments on two multi-modal classification benchmarks demonstrate significant robustness gains in various missing modality conditions. The first method shows more robust behavior under minimal missing conditions, where a single modality is absent, whereas the second performs better under extreme missing conditions, where all-but-one modalities are missing. Our code is available at this https URL.

[CV-11] MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

链接: https://arxiv.org/abs/2607.24665
作者: Yanhao Jia,Jiepeng Wang,Haibin Huang,Chi Zhang,Erik Cambria,Xuelong Li
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Machine Learning (cs.LG)
备注: 13 pages, 5 figures

点击查看摘要

Abstract:Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly enlarge total parameter counts and sparsity ratios without importing the efficiency mechanisms that made LLM scaling practical, so generation quality is seldom balanced against training and deployment cost. This raises a natural question: can the architectural principles behind efficient LLM scaling be adapted to AFMs in a more balanced way? We introduce ModernMOE (MMOE), a modernization of SiT-style diffusion transformers that systematically adapts routed experts, shared and lightweight experts, gate-residual routing, and attention-residual information reuse to AIGC generation. Rather than treating MoE as a single plug-in replacement, MMOE studies how different modern expert components affect convergence, efficiency, and generation quality when composed inside a diffusion transformer. Every experiment in this paper is trained on a single eight-GPU H100 node with batch size 256 for 400k steps, an accessible single-machine budget. Under matched training and sampling protocols and at this budget, MMOE reaches lower FID at every recorded checkpoint, that is, it converges faster per training step, than dense and intermediate sparse-expert baselines, and among the sparse variants it attains the best quality-cost balance. Routing analysis further shows stable expert specialization across depth, substantial use of lightweight routes, and modest step-to-step routing changes during denoising. These results suggest that AFMs can follow the balanced scaling path of LLMs by importing proven efficiency designs, rather than by simply increasing total parameters and sparsity ratios.

[CV-12] st-Time Adaptation via Dual Distillation for Videos Under Severe Distribution Shifts

链接: https://arxiv.org/abs/2607.24611
作者: André Sacilotti,Samuel Felipe dos Santos,Jurandy Almeida
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deep learning models have achieved state-of-the-art performance in several computer vision tasks. However, they experience severe performance degradation when applied to real-world scenarios due to unanticipated distribution shifts. Test-Time Adaptation (TTA) attempts to solve this problem by using unlabeled data from the target domain to dynamically adapt to the test distribution at inference time, without access to the source data. However, TTA remains a challenging problem when adapting to continuous, temporally correlated data, such as videos, and in scenarios where the target domain contains severe domain shifts. For this reason, few works in the literature explore TTA for videos under such extreme conditions. To overcome these limitations, we propose Test-time Adaptation via Dual Distillation (TADD), an online adaptation framework that relies on a lightweight projection adapter to bridge the domain gap. The adapter module is pre-trained on the source domain and then adapted to the target using our proposed complementary losses: (i) zero-shot distillation, which encourages alignment with the domain-agnostic features from a pre-trained vision-language model (VLM); and (ii) target distillation, which retains the source domain discriminative knowledge encoded in the pre-trained adapter. Built upon a frozen CLIP backbone, our method introduces this lightweight projection adapter as the sole updatable component during inference. We conducted extensive evaluations on three well-known video action recognition benchmarks: UCF-HMDB, Daily-DA, and Sports-DA. Our experiments in the closed-set scenario demonstrate that our method consistently outperforms state-of-the-art TTA baselines. Notably, our TTA approach improves upon previous methods by up to +3.81% on UCF-HMDB, +2.63% on Daily-DA, and +3.03% on Sports-DA.

[CV-13] QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

链接: https://arxiv.org/abs/2607.24598
作者: Arian Kheirandish,Fardin Ayar,Ehsan Javanmardi,Manabu Tsukada,Mahdi Javanmardi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project website: this https URL

点击查看摘要

Abstract:Video instance segmentation (VIS) requires models to detect, segment, and track object identities across frames, and most methods enforce temporal consistency through video-level supervision. Image-only training approaches, with MinVIS as one prominent example, have challenged this assumption, reaching competitive VIS without video training by treating frames as independent images and associating instances only at inference. The field has nonetheless moved toward ever more elaborate video-trained trackers, which depend on costly identity-consistent annotations, leaving the image-only direction under-explored. A diagnostic analysis identifies object query quality as the bottleneck: queries trained only to localize objects within a frame drift apart across frames and destabilize association. QueenVIS introduces a query-centric framework for strengthening image-trained VIS. During single-frame training, we enrich Mask2Former queries with two auxiliary heads: a feature-prediction loss that aligns each query with the pooled backbone descriptor of its instance, and a center-prediction loss that injects spatial structure. Both heads are discarded at inference, adding zero parameters, and temporal identity is maintained by a training-free query-propagation and memory-bank scheme. On YouTube-VIS and OVIS with a ResNet-50 backbone, QueenVIS improves over MinVIS, up to +6.7 AP on YouTube-VIS, +4.8 AP on OVIS, and +10.3 AP on the long-sequence YouTube-VIS split. QueenVIS achieves 50.9 AP on YouTube-VIS and remains competitive with recent video-supervised state-of-the-art, without processing a single video clip during training. Our findings suggest that strengthening the discriminative power and temporal stability of object queries is an important, underexplored axis for VIS. Code and models: this https URL

[CV-14] CameraAnything: Refilming Videos with Arbitrary Camera Control

链接: https://arxiv.org/abs/2607.24591
作者: Yixuan Li,Yanhong Zeng,Ka Leong Cheng,Jiayi Zhu,Hanlin Wang,Wen Wang,Yihao Meng,Hao Ouyang,Qiuyu Wang,Yue Yu,ZiDong Wang,Yiyuan Zhang,Yujun Shen,Dahua Lin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters. Existing approaches either rely on expensive 3D reconstruction to achieve full camera functionality or restrict editing to extrinsic parameter manipulation. Moreover, the coupled influence of intrinsic and extrinsic parameters on video appearance makes disentangled modeling particularly challenging. To address this, we adopt per-pixel Plücker ray injection alongside resolution-aware 3D RoPE in self-attention, building both camera conditioning and spatial positional encoding on the target latent to jointly control camera position, focal length, and native resolution editing without cropping or outpainting. To overcome the scarcity of paired training data, we further develop a scalable synthetic pipeline that constructs diverse dynamic scenes through structured multi-camera recording and generates synchronized videos with varied camera configurations. With a tailored orthogonal training strategy, CameraAnything enables expressive video reshooting with arbitrary viewpoint control, focal length adjustment, resolution adaptation, and multi-shot transitions within a single generation process, offering strong practical value for cinematic video editing and cross-platform content adaptation in video production.

[CV-15] CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

链接: https://arxiv.org/abs/2607.24582
作者: Jinlong Yang,Wenhao Zhang,Kuanwei Lin,Sijie Cheng
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example regardless of difficulty. This uniform strategy invokes unnecessary tool-assisted processing for easy questions and provides limited control when difficult questions require fine-grained temporal evidence. We propose CADER (Confidence-Aware Dynamic Evidence Reasoning), a training-free framework for adaptive and reliable long-video reasoning. CADER first performs global reasoning over uniformly sampled frames and estimates answer confidence with a logit-margin signal, allowing high-confidence examples to exit early. For uncertain examples, CADER activates a second-stage tool-augmented loop that combines temporal cropping, lightweight semantic verification, and Relevance-Guided Resampling to progressively localize question-relevant evidence. This design treats tool use as a sample-level decision: a single global pass handles easy cases, while additional reasoning is reserved for examples where uncertainty suggests that more evidence is needed. Experiments on multiple VideoQA benchmarks show that CADER improves long-video reasoning while bypassing Stage~2 for high-confidence samples. Moreover, when applied to a backbone trained only with tool-free chain-of-thought supervision, CADER achieves competitive performance against specialized tool-augmented frameworks, suggesting a practical inference-time route for adaptive long-video reasoning.

[CV-16] Denoising 3D images: robustness of persistent homology measures

链接: https://arxiv.org/abs/2607.24579
作者: Ebru Dagdelen,Aakash Karlekar,Manav Arora,Matthew Illingsworth,Jonathan Jaquette,Linda J. Cummings,Lou Kondic
类目: Computational Geometry (cs.CG); Computer Vision and Pattern Recognition (cs.CV); Algebraic Topology (math.AT)
备注: 25 pages

点击查看摘要

Abstract:When computing sub/super-level-set persistent homology (PH), the effect of noise may introduce millions of (short-lived) topological generators, presenting an obstacle to both the computation of PH of large 3D images, and any analysis of PH that incorporates the number of generators. As such, it is often necessary to denoise the data before computing its PH. We analyze the PH of synthetic 3D images of porous media in the presence of spatially uncorrelated noise, and perform a comparative analysis of various topological measures (e.g. bottleneck distance, Wasserstein distance, persistence statistics and persistence images) to assess their robustness to both noise and the denoising process (i.e. adding spatially uncorrelated Gaussian noise, and denoising by either a Gaussian convolution or a machine learning approach). Comments: 25 pages Subjects: Computational Geometry (cs.CG); Computer Vision and Pattern Recognition (cs.CV); Algebraic Topology (math.AT) Cite as: arXiv:2607.24579 [cs.CG] (or arXiv:2607.24579v1 [cs.CG] for this version) https://doi.org/10.48550/arXiv.2607.24579 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-17] he Visual Bottleneck: Sparse-Frame Adaptation of MLLM s for Joint Spatial-Temporal Video Grounding

链接: https://arxiv.org/abs/2607.24570
作者: Jiameng Zhang,Srikanth Madikeri
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 15 pages, 5 figures

点击查看摘要

Abstract:Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing every frame is infeasible at scale, so systems are constrained to sparse inputs of 8 to 16 frames per video. Yet state-of-the-art multimodal large language models (MLLMs) are pretrained on dense sequences of hundreds of frames, creating a fundamental mismatch between training and deployment conditions. This mismatch causes severe performance collapse: the Qwen3-VL 8B model drops from 56.0% to 22.3% temporal mIoU when frames are reduced to 16, a 60.2% relative degradation. We present a systematic empirical study of training strategies to close this gap for spatial-temporal video grounding. Our results suggest that visual feature extraction is the dominant bottleneck under sparse-frame inputs. Adapting only the final three ViT layers, 4% of total parameters, achieves 68.8% temporal mIoU and surpasses a zero-shot 8B model using dense inputs by 12.8 points. Language model fine-tuning, by contrast, offers negligible or negative returns. A boundary-aware sampling strategy, Hybrid16, further improves temporal mIoU by 26 points over uniform sampling when temporal boundaries are available. We conclude that for sparse-frame video grounding, training strategy dominates model scale: a fine-tuned 2B model consistently outperforms a zero-shot 8B model, with or without dense frame access. Comments: 15 pages, 5 figures Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.24570 [cs.CV] (or arXiv:2607.24570v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2607.24570 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[CV-18] EgoPlay: Event-Triggered Video Editing for Egocentric Streams SIGGRAPH

链接: https://arxiv.org/abs/2607.24560
作者: Jinjie Mai,Gordon Guocheng Qian,Willi Menapace,Arpit Sahni,Chaoyang Wang,Ashkan Mirzaei,Runjia Li,Sergey Tulyakov,Bernard Ghanem,Peter Wonka,Rameen Abdal
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted to SIGGRAPH Asia 2026 as a Conference Paper. Project page: this https URL

点击查看摘要

Abstract:We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form “when X happens, do Y,” EgoPlay infers whether and when event X occurs, preserves pre-event frames, and applies edit Y only to the post-event continuation. Rather than cascading a separate event detector with an editor, EgoPlay learns event recognition, temporal restraint, and pixel-level editing jointly in a single end-to-end model, while also handling negative and multi-event prompts. To support this, we construct a large-scale dataset of 106K event-triggered clip-prompt pairs spanning positive triggers, fabricated-trigger negatives, and multi-event prompts. We then train a bidirectional video diffusion editor with event-triggered supervision and derive a causal variant for chunk-by-chunk streamable inference. We further introduce an event-aware evaluation protocol that separately measures post-trigger editing quality, pre-trigger preservation, and false-trigger robustness. On the Ego4D benchmark, EgoPlay substantially outperforms EgoEdit, the state-of-the-art instruction-based egocentric video editing baseline, with relative gains of 17.7%, 16.9%, and 16.4% in editing quality, visual quality, and background consistency. It also surpasses a VLM-guided detector-editor baseline by 15.7%, 14.5%, and 13.5% on the same metrics, while using less than half the GPU memory.

[CV-19] FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

链接: https://arxiv.org/abs/2607.24522
作者: Kaiyang Ye,Yuan Ge,Junxiang Zhang,Bei Li,Ziming Zhu,Haishu Zhao,Xiaoqian Liu,Chenglong Wang,Jingbo Zhu,Zhengtao Yu,Tong Xiao
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state. Using the integral relation between trajectories and velocity fields, we derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps. Under a multi-reference setup, single-state FlowCTS-OPD outperforms vanilla KL-based OPD with faster convergence. FlowCTS-OPD improves GenEval from 0.90 to 0.93, OCR from 0.90 to 0.92, and PickScore from 22.75 to 23.06, while outperforming a mixed-reward RL baseline across all target metrics. Further analysis reveals a clear temporal supervision mismatch in vanilla KL-based OPD arising from its auxiliary SDE transition kernels. Beyond on-policy setting,FlowCTS also consistently outperforms vanilla SFT , particularly on OCR, while increasing supervision steps exhibit a trade-off between richer trajectory information and greater optimization difficulty.

[CV-20] DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

链接: https://arxiv.org/abs/2607.24516
作者: Jiahao Xie,Zhongbin Guo,Qianle Wang,Ruiqi Lu,Dongling Xiao,Wanxuan Sun,Cheng Yang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a principled, attributable criterion for admitting new data, while frontier recipes remain undisclosed. We formulate data construction as a systematic mixture-optimization problem and turn it into a reproducible engineering discipline by decoupling the mixture into two orthogonal sub-problems: inter-class ratios across capabilities and intra-class ratios within a category. For inter-class allocation, we use a single-variable iterative search; for intra-class composition, we apply a multidimensional, dataset-level assessment scoring Quality and Difficulty, and formulate selection as a constrained convex optimization with a diversity objective. The DecoupleMix framework delivers two critical capabilities: guiding what data to collect next and rendering dataset validation a controlled, attributable experiment. Experiments show our approach consistently surpasses heuristic baselines. Moreover, optimal ratios discovered on small-scale proxies transfer seamlessly to larger scales without retuning. Using 80B additional multimodal continue-pretraining tokens, our VLM is competitive with strong open-source models trained with substantially larger multimodal budgets.

[CV-21] NSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and Reconstruction

链接: https://arxiv.org/abs/2607.24495
作者: Jiaheng Li,Binsheng Zhang,Xinhai Chang,Wenzheng Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages, 9 figures

点击查看摘要

Abstract:Structured-light (SL) cameras power depth sensing in millions of devices, and recent neural SL decoding methods have substantially improved their depth quality. SLAM systems can benefit greatly from such strong depth sensing, where reliable geometry enables stable tracking and faithful reconstruction. In this work, we present NSL-SLAM, a practical SLAM system tailored for high-fidelity structured-light depth. We first strengthen SL depth sensing: inspired by the neural structured-light (NSL) method, we further incorporate strong monocular depth priors into the SL stereo decoding, reducing depth RMSE by 35% on Replica-SL compared to NSL. We then build a depth-centric SLAM pipeline with this stronger depth: because structured-light geometry is dense and metrically accurate, we keep it as the primary tracking signal, and add only sparse visual correspondences for geometrically degenerate cases and lightweight bundle adjustment for long-range drift. Our depth estimator and SLAM design reinforce each other: stronger depth makes a simple SLAM pipeline effective, and the depth-centric pipeline ensures this advantage transfers to downstream reconstruction. Experimentally, on the synthetic Replica-SL benchmark, NSL-SLAM achieves the best tracking accuracy and improves reconstruction F-score by 1.6 points over the SOTA baseline under a shared-depth protocol. On a real benchmark of 8 challenging scenes, it is the only method that avoids catastrophic failure on all sequences while achieving 43.3% lower trajectory deviation than selected baselines. The SLAM system runs online at 20.9 FPS, demonstrating that stronger structured-light depth and depth-centric system design together enable practical, robust SLAM.

[CV-22] Rethinking Expert Training for Model Merging with Prompt Learning

链接: https://arxiv.org/abs/2607.24465
作者: Christos Georgakilas,Aniello Panariello,Samir El Karrat Moreno,Simone Calderara,Dimosthenis Karatzas,Joost van de Weijer
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 14 pages

点击查看摘要

Abstract:Model merging aims to combine multiple domain-specialized experts trained from a shared foundation model into a single multi-task model. Existing approaches largely focus on improving the merging procedure itself and typically assume experts obtained through full-parameter fine-tuning. In this work, we revisit expert training for model merging. We first show that prompt-based adaptation provides a strong baseline: independently learned prompts can be exploited across tasks while keeping the backbone fixed, avoiding the interference introduced by weight merging. Building on this observation, we introduce Dual-Tuned Experts (DTEs), a two-stage training strategy that first learns prompts and then fine-tunes the vision encoder. This reduces the magnitude of task-specific parameter updates and produces experts with higher merge compatibility. Experiments across multiple CLIP architectures, full fine-tuning, and LoRA experts show that DTEs consistently improve merged performance of standard merging approaches and remain effective even when combining heterogeneous sets of experts.

[CV-23] ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image

链接: https://arxiv.org/abs/2607.24453
作者: Mingzhi Xu,Yizhe Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 15 pages, 2 this http URL by PRCV 2026

点击查看摘要

Abstract:Learning from minimal human supervision is a long-standing goal in medical image analysis, where dense expert annotations are costly. We study retinal vessel segmentation in an extreme semi-supervised setting with one annotated image and a pool of unlabeled images. We propose ESRVS, which selects a representative reference image for manual annotation and transfers vessel cues using target-domain-adapted DINOv3 features. ESRVS constructs a multi granular vessel prototype, combines prototype-similarity maps with a physics-inspired prior to generate initial pseudo-labels, and refines the transferred supervision through weighted pseudo-label training and adversarial refinement. Across eight public datasets, ESRVS achieves the best Dice and clDice on six datasets, and the best HD95 on all eight datasets among the compared semi-supervised methods, although those methods use 10 to 20% labeled data. With Mask2Former, ESRVS retains on average 93.7% of fully supervised Dice and 95.1% of fully supervised clDice. These results demonstrate the potential of foundation-model label propagation for highly label-efficient retinal vessel segmentation. Code is available at this https URL.

[CV-24] RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models

链接: https://arxiv.org/abs/2607.24447
作者: Qihui Zhu,Yuchen Wang,Zijian Wen,Tao Zhang,Mengjie Zhang,Yang Liu,Shuangwu Chen,Siying Wu,Jian Yang,Xiaofeng Jiang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories generated by the student. However, existing methods often rely on verified solution traces, explanations generated by external models, or manually localized visual evidence, which limits their scalable application to multimodal large language models. To address this issue, we exploit the information gap between high- and low-resolution views of the same image and propose RP-OPSD (Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models). During training, the student policy generates on-policy trajectories from images at one-quarter of the original resolution, while the teacher policy provides supervision using the original-resolution images. By minimizing the divergence between their output distributions along the student trajectories, the student learns the predictive behavior of the teacher under high-resolution inputs, thereby strengthening its low-resolution capability and transferring the learned improvement to original-resolution inference. RP-OPSD requires neither additional human annotations nor external models to generate solution traces but only image–question pairs. Experiments on Qwen3.5-9B show that RP-OPSD achieves a 5.45% relative improvement in average performance at the original resolution and a 1.78\times training speedup over OPSD. These results demonstrate that resolution differences can serve as a simple and scalable source of privileged information, providing an effective and efficient approach to on-policy self-distillation for multimodal large language models.

[CV-25] MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

链接: https://arxiv.org/abs/2607.24436
作者: Dehao Hao,Kaiyi Zhang,Tanghui Jia,Xiangjun Gao,Dongyu Yan,Weikai Chen,Zeyu Hu,Lingting Zhu,Yingda Yin,Runze Zhang,Li Yuan,Xin Wang,Long Quan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. Existing approaches largely follow two paradigms: sparse voxel-based representations achieve strong reconstruction quality but incur significant memory and computational overhead, while set-based representations are compact and continuous yet typically lag in fidelity due to latent sparsity and excessive global smoothness. We propose MSVS-VAE, a hierarchical set-based VAE that closes this fidelity gap without sacrificing compactness. Our key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling. To efficiently decode from the densified hierarchy, we replace global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the exhaustive latent set. We further introduce multi-scale query decoding to fuse coarse-to-fine latent features, where coarse scales provide stable global context, and fine scales refine localized geometry, reducing artifacts from overly local receptive fields. Extensive experiments on Objaverse, ABO, and in-the-wild benchmarks demonstrate that MSVS-VAE consistently outperforms prior set-based and voxel-based VAEs, delivering approximately 10x faster decoding than prior set-based methods and approximately 10x higher compactness than voxel-based baselines.

[CV-26] InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting

链接: https://arxiv.org/abs/2607.24431
作者: Qi Zhang,Xinquan Yu,Kaiyi Zhang,Hui Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages, 7 figures

点击查看摘要

Abstract:Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is critical for driving safety. Even though current methods have achieved good performance, the strong spatial-temporal modeling between the input multi-view frames is still underexplored, which limits the performance of those methods in future 4D forecasting. To address this gap, we introduce a novel framework, InterOCF, for 4D occupancy forecasting that jointly models temporal dynamics in both 3D voxel-based representations and multi-view segmentation sequences, while explicitly incorporating feature interaction between the 2D and 3D branches. Our framework incorporates three core components: 1) A 3D Spatio-Temporal (3DST) module that learns volumetric dynamics from historical voxel states to predict future voxel states; 2) A 2D Spatio-Temporal (2DST) module employing an auxiliary multi-view temporal segmentation forecasting task to enhance temporal semantic dynamics; 3) A Spatio-Temporal Interaction Modeling (STIM) module that enables feature interaction between 2D and 3D representations. Experiments on the nuScenes, Lyft-Level5, and nuScenes-Occupancy datasets show that InterOCF consistently outperforms existing baseline approaches.

[CV-27] MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

链接: https://arxiv.org/abs/2607.24424
作者: Shaofei Lei
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages

点击查看摘要

Abstract:Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency. We introduce \method, a Multi-scale Adaptive Vision Encoder. \method uses position-dependent gates to fuse shallow, intermediate, and deep features from a vision Transformer, preserving global semantics while enhancing edges, text, and local structure. It then performs question-conditioned token routing according to question relevance, local information content, global semantics, and spatial coverage, with a token budget that adapts to image complexity. To mitigate compression loss, we further introduce full-to-compressed representation distillation and a spatial diversity regularizer. In an illustrative simulation under a unified 7B language-model framework, \method reduces the average number of SigLIP-SO400M visual tokens from 729 to 146 (approximately 80.0%) and improves the mean score on VQAv2, GQA, TextVQA, ScienceQA-IMG, and MMBench by 2.2 percentage points, while reducing single-image time to first token from 228,ms to 129,ms. We provide the full model design and evaluation protocol. All reported numbers currently serve only as placeholders for paper organization and experimental design; formal claims require real training runs, independent replications, and official benchmark evaluation.

[CV-28] IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data

链接: https://arxiv.org/abs/2607.24422
作者: Tahar Chettaoui,Guray Ozgur,Eduarda Caldeira,Arturas Nakvosas,Hatef Otroshi Shahreza,Sébastien Marcel,Rishabh Shukla,Aditya Takkar,Rushil Khullar,Lalak Yadav,Gourav Gupta,Anant Gupta,Shiqi Yu,Vitomir Struc,Naser Damer,Fadi Boutros
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the IEEE International Joint Conference on Biometrics 2026 (IJCB 2026)

点击查看摘要

Abstract:This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2026 International Joint Conference on Biometrics (IJCB 2026). The competition received a total of eight valid submissions from four distinct teams across two complementary tracks: a Full Data Track, in which participants adapt the CLIP ViT-L/14 foundation model using large-scale synthetic identity data, and a Limited Data Track, designed to reflect more resource-constrained adaptation regimes. All training data was generated exclusively using IDPERTURB. Submitted solutions are ranked based on verification and identification performance across a diverse suite of benchmarks, including LFW, CFP-FP, AgeDB-30, CALFW, CPLFW, IJB-B, IJB-C, and TinyFace, using the Borda count method. Fairness evaluation is additionally conducted on the RFW dataset across four demographic groups. The results demonstrate that adaptation of the CLIP foundation model with synthetic training data substantially improves over the off-the-shelf model and, in several cases, surpasses the baseline. Notably, full fine-tuning with Sub-Center ArcFace (DMSTI-Neurotechnology) leads the Full Data Track, while rank-stabilized LoRA adaptation (Idiap-BSP) proves most effective under limited-data conditions.

[CV-29] Accuracy potential of visual localization exploiting high-end street-level imagery

链接: https://arxiv.org/abs/2607.24409
作者: Jonas Meyer,Stephan Nebiker,Pascal Theiler,Norbert Haala
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 26 pages, 6 figures

点击查看摘要

Abstract:Accurate and reliable pose information with respect to a reference frame is increasingly demanded across applications such as autonomous navigation, surveying, robotics, and augmented and mixed reality. Visual localization can serve as a complementary positioning modality to GNSS, whose applicability and accuracy are often limited. Yet, the accuracy potential of visual localization has not been systematically investigated against survey-grade demands. This is mainly due to the lack of publicly available, large-scale outdoor datasets with ground-truth poses in the sub-centimeter range. In this work, we address both gaps. We introduce a scalable visual localization pipeline that employs precisely georeferenced, high-resolution street-level imagery directly as the scene representation. It combines prior-guided reference candidate selection with on-the-fly local Structure-from-Motion reconstruction and PnP-based pose estimation. We further present the FHNW Muttenz dataset, a real-world dataset covering a contiguous 10 km street network mapped in two mobile mapping campaigns approximately 1.5 years apart. It consists of high-resolution reference imagery and query sequences acquired by four different cameras across five representative scenes. All images are precisely co-registered, yielding 6-DoF ground-truth poses in the sub-centimeter range. Using this dataset, we evaluate the accuracy potential of visual localization. Our experiments demonstrate median pose accuracies in the range of 1-5 cm for translation and 0.05-0.1° for rotation, reaching as low as 1 cm and 0.03° under favorable conditions. These results show that visual localization can complement survey-grade GNSS positioning, paving the way for 3D geospatial data acquisition using consumer devices and fully automated georeferencing approaches. The dataset is publicly available at: this https URL.

[CV-30] Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding ACM-MM2026

链接: https://arxiv.org/abs/2607.24407
作者: Tianyi Gao,Han Fang,Tianyi Ding,Hao Li,Xin Wei,Hongbo Sun,Xiaodong Dong,Ye Yuan,Jinglin Xu,Kongming Liang,Hao Sun,Jingmin Xin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ACM MM 2026

点击查看摘要

Abstract:Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing, text-based methods rely on coordinates or index prediction, severely limiting the perceptual capabilities of the model for dense visual objects. Meanwhile, latent token-based methods employ special tokens without inherent spatial references and use a decoding mechanism that lacks thinking steps, weakening high-level reasoning capabilities. Consequently, developing a unified framework that excels in both perception and reasoning remains challenging. To address this, we propose Mixture-of-Thought-Tokens (Motto), a new free-form multimodal grounding method that bridges the perception-reasoning gap, enabling MLLMs to empower diverse, arbitrary grounding queries. Specifically, we introduce Spatially-Grounded Thought Tokenization to explicitly align special tokens with spatial locations for clear spatial correspondence and visual interpretability. We further design a Context-Adaptive Chain-of-Tokens that dynamically switch grounding modes within an interleaved reasoning chain, achieving robust grounding across tasks of varying complexity. In addition, we construct PR-Bench, a new referring expression comprehension benchmark to evaluate the perception-reasoning gap. Extensive experiments demonstrate that Motto achieves state-of-the-art performance across diverse free-form grounding tasks.

[CV-31] GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion

链接: https://arxiv.org/abs/2607.24403
作者: Qiang Hu,Zhenlong Wu,Lei Huang,Zihan Zheng,Xiaoyun Zhang,Wenjun Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are costly to store and transmit. Existing feed-forward Gaussian compression methods formulate decoding as deterministic representation recovery, which becomes inadequate at low bitrates when high-frequency textures and view-dependent appearance are discarded. Although generative models offer a promising alternative, using them as standalone post-processing decouples generation from the transmitted scene structure, thereby compromising cross-view consistency. To address these limitations, we propose GenSplatCodec, a unified feed-forward Gaussian codec that reformulates low-bitrate Gaussian compression as geometry-guided generative decoding. We present a detail-aware feed-forward Gaussian coding scheme within a dual-stream formulation, where the resulting compact Gaussian structural stream is complemented by a lightweight reference appearance stream. We further introduce a geometry-guided one-step generative decoding approach that jointly exploits decoded structural and appearance cues through hierarchical geometry control to reconstruct high-fidelity and view-consistent novel views. Finally, we develop a three-stage optimization strategy that stabilizes the learning of the unified codec and adapts the generative decoder to codec-derived structural and appearance cues. Extensive experiments across multiple datasets demonstrate that GenSplatCodec consistently achieves superior rate-distortion (RD) performance over existing methods.

[CV-32] Ambient pressure compensation and robust position control of oil-filled electric joint systems for underwater manipulators

链接: https://arxiv.org/abs/2607.24384
作者: Hongrui Wu,Songhui Wang,Xin Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Electric joint systems are significant elements of an underwater manipulator for its actuation, drive, and control. Working in an underwater environment, the joints suffer huge ambient pressure. To withstand it, the pressure compensation method is usually deployed, whereas the pressurized oil introduces sealing problems as well as parametric uncertainties and unknown disturbances for the dynamic model of the joint. To tackle these issues, this study proposes a design framework for the underwater oil-filled electric joint. The dynamics of the pressure compensation module is analyzed and the structure of the joint is optimized to seal the internal hydraulic oil. An uncertainty dynamic model of the oil-filled joint is established and a robust position controller is designed based on the structured singular value synthesis (mu-synthesis). Experimental results validate the feasibility of the proposed methods.

[CV-33] MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention

链接: https://arxiv.org/abs/2607.24377
作者: Jianlin Yu,Jing Lin,Linghui Kong,Aiyue Chen,Weiyi Sun,Chenyu Zeng,Wangli Lan,Jinxi Li,Zhuo Zheng,Ziyang Yue,Danning Ke,Fei Yi,Tianchi Hu,Yuan Ding,Yiwu Yao,Junsong Wang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling and the row-wise normalization error introduced in the softmax loop. We propose MXAttention, a data-free post-training quantization framework for MXFP4 attention. MXAttention introduces two components: Universal Optimal Scaling (UOS), which exploits the periodic structure of power-of-two microscaling to derive a distribution-independent optimal scaling boundary Qmax=7.25 without calibration or search, and Pre-Normalization Quantization (PNQ), which quantizes unnormalized softmax exponentials before row-wise summation to preserve normalization by construction. Experiments on Wan2.2 and HunyuanVideo show that MXAttention closes at least 95% of the VBench Imaging Quality gap between OCP MXFP4 and FP16, substantially improves frame-level similarity, and preserves FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics. MXAttention also achieves performance competitive with strong NVFP4-based baselines with negligible overhead when fused into the attention pipeline. The implementation is publicly available in MindIE-SD.

[CV-34] HistoGPA: A Context-Conditioned Gene-Prior Attention Framework for Histology-Based Spatial Gene Expression Prediction

链接: https://arxiv.org/abs/2607.24364
作者: Ziang Liu,Xinhai Chen,Yigui Feng,Shuai Li,Qingyang Zhang,Jie Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 2 figures, 3 tables

点击查看摘要

Abstract:Predicting spatial gene expression from routine hematoxylin and eosin (HE) images provides a practical complement to experimental spatial transcriptomics. Existing approaches focus on local or multi-scale visual features and often treat pretrained gene representations as fixed priors, although the interpretation of local morphology and the relevance of gene priors depend on tissue context. We propose HistoGPA, a context-conditioned gene-prior attention framework that uses a shared slide-level representation in two parallel pathways: one modulates local morphological features, whereas the other conditions pretrained gene embeddings and retrieves gene-prior information through cross-attention. This design enables each spatial location to retrieve context-adapted gene-prior information using its local morphology, position, and slide context. Across ten cancer types in HEST-1k, HistoGPA achieves the highest macro-averaged gene-wise Pearson correlation coefficient among the compared methods under the same evaluation protocol for both the top-50 and top-1,500 highly variable gene sets. Additional analyses show that HistoGPA better recovers the spatial expression patterns of cancer-associated genes and yields greater agreement between clusters derived independently from predicted and ground-truth expression profiles. Together, these findings motivate a context-dependent view of histology-to-expression prediction, in which local morphological representations and gene priors are jointly adapted to the broader tissue context.

[CV-35] aoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

链接: https://arxiv.org/abs/2607.24359
作者: Qijun Gan,Chenwei Zhang,Meiguang Jin,Junfeng Ma,Qiu Shen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video synchronization across successive segments. A bounded cache retains local motion and phonetic context but discards older evidence, whereas attending to the complete generated history is computationally expensive and can propagate accumulated errors. We present \method, an anchor-guided persistent-memory framework for few-step joint audio-video generation. The framework preserves an immutable visual anchor, compresses completed video and audio blocks into fixed-capacity dynamic states, and retrieves those states through modality-specific residual attention without extending the active cache. A reference-aware modulation method additionally conditions video features on dynamic and anchor appearance statistics. Anchor-preserving causal-context distillation varies rollout horizon, prefix provenance, and cache-history reliability while keeping the immutable visual anchor unperturbed. By separating persistent memory from stage-local denoising dependencies, \method further admits stage-parallel execution across blocks, accelerating autoregressive inference without pipeline-specific retraining. We evaluate long-form video continuations with appearance, temporal, synchronization, facial, and speech diagnostics. Results show that \method preserves stable appearance across prompt-conditioned segments and strong audio-visual synchronization under autoregressive generation. Our project page is this https URL.

[CV-36] PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation

链接: https://arxiv.org/abs/2607.24353
作者: Guo Tang,HongJie Luo,Tianxu Wang,Ying Zhang,Hao Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 5 figures

点击查看摘要

Abstract:Text-to-image generation models can synthesize high-quality images from natural language descriptions, but their performance remains highly sensitive to prompt formulation. Existing prompt optimization methods mainly rely on text-side rewriting, prompt expansion, or external reward signals, offering limited image-grounded diagnosis and weak support for learning reusable optimisation policies. In this paper, we propose PRISM, a Prompt Refinement framework via Image-grounded Self-rewarding Mechanism. PRISM closes the prompt-image-feedback loop by interpreting generated images with structured visual diagnosis and scoring them along semantic consistency, aesthetic quality, and human preference alignment. It first initializes a unified VLM through multi-task supervised fine-tuning, and then improves the prompt policy via self-rewarding optimization with a hybrid ideal-point and Chebyshev reward. Extensive experiments show that PRISM improves holistic image quality and fine-grained semantic alignment, while providing interpretable feedback for targeted prompt refinement. The code is available at this https URL.

[CV-37] Multiview Multi-Person Human Mesh Recovery Under Large Scenes with Occlusions

链接: https://arxiv.org/abs/2607.24302
作者: Qi Zhang,Tao Yu,Jiechao He,Antoni B. Chan,Hui Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 7 figures

点击查看摘要

Abstract:Human mesh recovery (HMR) aims to recover 3D human meshes from images. Most existing HMR benchmarks and methods focus on either multi-person reconstruction from a single view or single-person reconstruction from multiple views, where the number of subjects and the scene scale are relatively limited. Such settings are insufficient for real-world applications with large scenes and severe inter-person occlusions. To address this limitation, we introduce a large-scale synthetic benchmark for multiview multi-person HMR, termed MVMP-HMR. The proposed dataset contains 15 complex scenes with up to 50 camera views and 30 interacting persons, featuring large spatial coverage and severe occlusions, which significantly increases the difficulty of human mesh recovery. Based on this benchmark, we further propose a multiview multi-person whole-body human mesh recovery model, referred to as MVMP-HMR model. The model first fuses multiview features into a scene-level 3D feature volume, and then leverages pelvis joints predicted by a 3D pose estimation network to extract person-specific queries from the 3D feature volume. These human queries are cross-attended with the 3D feature volume and integrated to decode each person’s 3D mesh. Moreover, we introduce two novel losses–the orientation loss and the 3D joint density loss–to alleviate orientation and pose ambiguities under severe occlusions. Experiments demonstrate that existing state-of-the-art HMR methods struggle on the proposed MVMP-HMR benchmark, while our method consistently outperforms prior SOTAs in large-scale scenes with severe occlusions.

[CV-38] UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing

链接: https://arxiv.org/abs/2607.24298
作者: Zefan Qu,Zhenwei Wang,Gerhard Petrus Hancke,Rynson W.H. Lau
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL

点击查看摘要

Abstract:Recent 3D foundation models can generate high-quality assets from a single image, but degrade markedly on unconstrained multi-image inputs, often producing distorted geometry, over-smoothed textures, and chaotic colors. We argue that this failure stems not from limited model capacity, but from a mismatch between single-image cross-attention and the multi-image setting: existing models lack a principled way to decide which image each 3D voxel should trust at each denoising step. Revisiting recent single-image 3D foundation models, we show that explicitly routing each voxel to its most informative image is sufficient to unlock strong performance on inconsistent multi-image inputs. Based on this observation, we propose UMI3D, a training-free and plug-and-play framework that restructures cross-attention for unconstrained multi-image 3D generation. Its core, Simultaneous Focus Cross-Attention (SFC-Attn), activates all conditioning images at each denoising step while allowing each voxel to focus on the single image that best explains it. To enable this routing, we derive the Voxel Reference Score (VRS), a model-intrinsic metric for voxel–image affinity that requires no external matching, segmentation, or correspondence models. Extensive experiments show that UMI3D unlocks the multi-image potential of single-image 3D generation frameworks across diverse tasks. Project Page: this http URL.

[CV-39] Superpixel-Based QUBO for Scalable Quantum-Enhanced Medical Image Segmentation

链接: https://arxiv.org/abs/2607.24288
作者: Mohammad Chalhoub,Mahdi Chehimi,Laia Domingo,Omar Alhussein,Ahmed Farouk,Saif Al-Kuwari
类目: Computer Vision and Pattern Recognition (cs.CV); Networking and Internet Architecture (cs.NI); Quantum Physics (quant-ph)
备注: 7 pages, 2 figures

点击查看摘要

Abstract:Quadratic unconstrained binary optimization (QUBO) has emerged as a powerful framework for medical computing problems. Binary decision variables naturally represent clinical choices, making QUBO formulations well-suited for quantum annealing hardware. However, a fundamental scalability challenge limits practical deployment: problem size grows rapidly with input dimensionality, creating computational bottlenecks that restrict applications to simplified scenarios. This paper addresses this challenge through hierarchical problem reduction, as demonstrated in medical image segmentation, where pixel-level QUBO formulations create over 65,000 variables for a 256x256 image, forcing existing approaches to downsample to 42x42 resolution and discard 97% of pixel information. A superpixel-based QUBO framework is proposed using simple linear iterative clustering (SLIC) to group pixels into perceptually meaningful regions, then formulate segmentation as QUBO over a region adjacency graph (RAG) combining min-cut and smoothness objectives. Validation on INbreast mammography breast cancer images demonstrates a 4.2% improvement in segmentation quality (mean IoU 0.76 vs 0.73) with 33 computational speedup (0.67s vs 21.97s) and a 97.3% reduction in problem size (1764 to 48 variables), all achieved while processing full-resolution images rather than downsampled versions. The reduced problem size also fits well within current quantum annealer connectivity limits, removing the embedding overhead that has historically blocked direct deployment of pixel-level QUBO segmentation on quantum hardware.

[CV-40] SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth Estimation IROS

链接: https://arxiv.org/abs/2607.24249
作者: Tarun R,Anuj Verma,Laksh Nanwani,Sourav Garg,K. Madhava Krishna
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026

点击查看摘要

Abstract:Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navigation hazards. While specialized hardware sensors can detect glass, they lack modularity and have extensive hardware dependencies. Consequently, learning-based monocular depth estimation has emerged as a compelling alternative. However, domain-specific glass-aware monocular depth estimators struggle with unfamiliar indoor layouts; restricted by the severe scarcity of real-world glass depth annotations, they fail to generalize zero-shot to new settings. This motivates us to explore whether the extensive priors of text-to-image diffusion models can enable generalizable perception of transparent surfaces. We introduce SILICA, a unified pipeline leveraging these priors to jointly predict glass segmentation and glass-aware depth. This mutual information exchange establishes a robust spatial hierarchy, entirely eliminating the need for paired real-world glass depth annotations. Subsequently, we use the predicted segmentation mask to explicitly filter incorrect glass depth points from standard sensors, recovering accurate metric glass depth for downstream 3D mapping and autonomous collision avoidance. Supported by our novel Mirage 18k dataset, extensive experiments demonstrate that SILICA achieves remarkable zero-shot transfer across diverse, unseen environments, outperforming state-of-the-art models by almost 20% and setting a new benchmark for transparent surface perception.

[CV-41] FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

链接: https://arxiv.org/abs/2607.24241
作者: Shengyi Wang,Niantong Li,Guangzheng Hu,Hong Qi,Fei Ding,Weixu Qiao,Jinlin Wang,Xiaotong Lv,Peng Han,Zimeng Li,Fanshu Ding,Yushu Wang,Han Wu,Jingjing Chen,Chongxiao Wang,Yanhao Wu,Chenglong Huang,Xiaoqian Zhu,Jie Tian,Hua Li,Jingjing Fan,Mingshuang Tang,Zhong Li,Hengxia Qiang,Weibin Chen,Jinyang Zhen,Bing Zhao,Lin Qu,Jing Li,Hu Wei
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \rho = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.

[CV-42] MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving

链接: https://arxiv.org/abs/2607.24224
作者: Junchen Huo,Wanming Hao,Song Wang,Enqing Chen,Shouyi Yang,Guanghui Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages

点击查看摘要

Abstract:Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a unified bird’s-eye-view (BEV) feature map for the joint learning of multiple perception tasks. However, such a single feature map hardly carries sufficient information to simultaneously meet the requirements of various perception tasks, leading to a very limited perception performance. To mitigate this limitation, this paper proposes MATS, a novel multi-modality multi-task learning approach with modality-adaptive BEV fusion and task-specific Mixture-of-Experts (MoE) for 3D perception. Specifically, a simple modality-adaptive BEV fusion module is designed to adaptively recalibrate the BEV features by modeling the global cross-modality dependencies, generating diverse BEV feature maps for various perception tasks. For joint multi-task learning, this paper proposes a task-specific MoE module to decouple the tasks and enable the network to automatically choose the appropriate BEV feature candidates for each specific task. To validate the effectiveness of the proposed approach, we conduct extensive experiments on the large-scale benchmark nuScenes. With the camera- and LiDAR-modality input data, the proposed approach outperforms the state-of-the-art (SOTA) by a significant margin. Furthermore, the experimental results on the single tasks show that the proposed approach significantly outperforms the baselines. The code and trained models will be available upon publication.

[CV-43] reeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation

链接: https://arxiv.org/abs/2607.24215
作者: Yuze Sun,Zhongjie Duan,Yingda Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Although general text-to-image models excel in open-domain generation, their performance degrades significantly in specialized downstream domains, particularly when generating images of rare biological species. Hindered by long-tailed distributions, general models struggle to capture subtle fine-grained details, while per-species fine-tuning methods over-isolate individual species and consequently ignore the shared visual features among closely related taxa. To address this, we propose TreeAdapter, a novel framework that explicitly leverages hierarchical taxonomic data. Rather than using a monolithic model or independent per-species modules, TreeAdapter attaches lightweight adapters to every node of the taxonomic tree. Specifically, leaf-node adapters capture species-specific visual traits, while internal-node adapters encapsulate shared semantics among descendant taxa. We introduce a two-stage training paradigm where ancestor adapters are optimized to model only the residual visual features unexplained by their descendants. This model architecture and training paradigm enable the model to fully leverage hierarchical information, ensuring the accurate generation of visual features for each species. Extensive experiments across three large-scale biodiversity benchmarks demonstrate that TreeAdapter achieves state-of-the-art fine-grained generation quality, outperforming both general-purpose and domain-specific baselines.

[CV-44] Effect of User-Prompted Priors on Semi-Automated Cancer Lesion Segmentation in Whole-Body Computed Tomography

链接: https://arxiv.org/abs/2607.24210
作者: Isac Stark,Johan Öfverstedt,Elin Lundström,Simon Ekström,Håkan Ahlström,Joel Kullberg
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at MIUA (Medical Image Understanding and Analysis). To appear in Lecture Notes in Computer Science (LNCS). 12 pages, 2 figures

点击查看摘要

Abstract:In clinical oncology studies, metastatic cancer is commonly evaluated using “Response Evaluation Criteria in Solid Tumors” (RECIST), in which the diameter of up to five lesions is measured and followed over the course of treatment. However, RECIST shows limited correlation with overall survival. Total tumour volume (TTV) is a stronger predictor but typically relies on manual ground-truth segmentation of all lesions, which is time-consuming and requires expert domain knowledge. Semi-automated approaches leveraging user-prompted priors, such as bounding boxes and single-slice contours, as inputs to automated segmentation methods can facilitate the generation of ground-truth segmentations. This work investigates the impact of different user-prompted priors on semi-automated cancer lesion segmentation performance in whole-body computed tomography. Across 3-fold cross-validation and external testing, more complex spatial priors consistently improved performance, with contour priors from three orthogonal planes (axial, coronal and sagittal) achieving the best results. On the external test (n=3865 lesions), this approach achieved a mean Dice score of 0.882, compared to a mean Dice score of 0.671 for the baseline model with no spatial prior. These findings suggest that the use of multi-plane orthogonal user-prompted priors can improve semi-automated tumour lesion segmentation and support efficient generation of high-quality volumetric ground-truth data.

[CV-45] Surgical Re-enactment for Operating Room Workflow Datasets ICRA2026

链接: https://arxiv.org/abs/2607.24206
作者: Jana Nina Friedrich,Andrea Karin Maria Ross,Angelo Henriques,Mario Peter Martin Weisser,Ling Zhang,Mohammad Ali Nasseri
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the 4th “Robot-Assisted Medical Imaging” (RAMI) workshop as part of ICRA 2026

点击查看摘要

Abstract:The introduction of new technologies, such as surgical robots, is driving the vision of a connected, smart operating room (OR). However, realizing this vision requires a deep understanding of surgical workflows, which relies on realistic datasets capturing the actions of all OR personnel from both full room and surgical field perspectives. Acquiring such data in real ORs is prohibitively challenging due to factors such as ethics committee approvals, limited space for camera installation, and sterility regulations preventing the use of tracking markers. We present a step-by-step methodology for re-enacting complete surgical procedures in a reconstructed OR. This approach enables the creation of repeatable and annotatable workflow datasets for training activity recognition models, generating scene graphs, and formalizing surgical process models. Developed for robot-assisted ophthalmic surgery, our methodology combines expert consultation, structured workflow formalization, OR reconstruction, role-based training, real OR observation, and iterative recording with post-take debriefing. We provide concrete recommendations to allow other research groups to seamlessly adopt this methodology for their own surgical domains.

[CV-46] Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

链接: https://arxiv.org/abs/2607.24199
作者: Yueru Luo,Xu Yan,Changqing Zhou,Yiming Yang,Chao Zhan,Shuqi Mei,Chao Zheng,Zhen Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Technical Report

点击查看摘要

Abstract:Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversity and context dependence of traffic signage. Importantly, regulation understanding is not a simple recognition task, but a reasoning problem: whether a rule applies depends on interpreting the sign in relation to the spatial layout of lanes and scene context. To support such reasoning, MapDR provide fine-grained annotations that link each traffic sign’s regulatory rules to the specific lanes they govern. Existing methods, however, largely treat this as direct sequence prediction, ignoring the underlying reasoning that connects sign semantics and map structure. To address this limitation, we explicitly incorporate reasoning into this task and propose a framework that equips vision-language models (VLMs) with chain-of-thought (CoT) capabilities. We first design a scalable CoT curation pipeline that bootstraps rationales from a strong LLM through a two-round strategy and employs a VLM-based verifier to filter out incorrect cases, yielding a high-quality set of (CoT, answer) pairs. Building on this foundation, we adopt a two-stage training scheme: supervised fine-tuning (SFT) to teach rationale-to-answer generation, followed by GRPO reinforcement learning with answer-grounded, fine-grained rewards to further improve final answer accuracy. Extensive experiments on MapDR show that our approach significantly improves both interpretability and accuracy, establishing the first reasoning-based framework for regulation-aware autonomous driving.

[CV-47] Face Age Verification Vulnerabilities Under Simple Appearance Manipulations

链接: https://arxiv.org/abs/2607.24194
作者: Ioannis Sarridis,Ioannis Kompatsiaris,Symeon Papadopoulos
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Online platforms increasingly rely on automated age estimation systems to enforce minimum-age policies. Focusing on vision-based models designed for this task, concerns arise regarding their robustness to simple appearance changes that underage individuals may use to bypass such systems, such as drawing a mustache or applying lipstick. In this work, we present a systematic study of age verification robustness by simulating visual alterations that can be easily achieved by underage individuals. We evaluate seven models, including vision, vision-language, and multimodal large language models, across three datasets and four manipulation types. Interestingly, under drawn beard stubble, up to 61% of True Negatives are flipped into False Positives. Furthermore, we investigate how different demographics are affected by such manipulations, finding that Indians are more affected by beard stubble manipulations, while females are more affected than males across all manipulations. Finally, we explore how these biases can be mitigated using bias mitigation methodologies in lightweight linear probe settings.

[CV-48] LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target Detection

链接: https://arxiv.org/abs/2607.24184
作者: Yuhao Fan,Le Hui,Yuchao Dai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak targets in infrared imagery usually occupy only a few pixels and are easily submerged by cloud clutter, ground edges, and bright noise, making it difficult for lightweight segmentation-based methods to preserve local target structures while suppressing background interference. To address these challenges, we propose LCMamNet, a lightweight cross-scale Mamba network that progressively enhances local target structures, interacts cross-scale context in a latent space, and restores spatial details with background suppression. Specifically, a compact hierarchical encoder with cross-shaped directional bottleneck residual (CDBR) blocks strengthens direction-sensitive target structures under a small computation budget. A latent dense cross-scale fusion (LDCF) module then performs dense all-level interaction through bidirectional Mamba modeling and reorganizes the interacted features into stable hierarchical semantics. Finally, a progressive decoder selectively recovers shallow spatial details while suppressing irrelevant background textures. Extensive experiments on IRSTD-1k, NUAA-SIRST, and NUDT-SIRST show that the proposed network achieves mIoU scores of 71.25%, 79.60%, and 95.58%, respectively, with only 1.175M parameters and 6.91 GFLOPs. It also runs with a mean inference latency of 6.62 ms, and deployment results on an NVIDIA Jetson Orin NX 16G SUPER further demonstrate its practical potential for real-time edge inference. The code and checkpoints are publicly available at this https URL.

[CV-49] DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

链接: https://arxiv.org/abs/2607.24159
作者: Mengqi Zhang,Sahil Khose,Simar Kareer,Yuchen Song,Unnat Jain,Judy Hoffman
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Project page with videos, code, and checkpoints: this https URL

点击查看摘要

Abstract:Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations. Video generative models offer a promising foundation by encoding rich spatiotemporal priors through future predictions. However, existing Video-Action Models either couple video and action prediction in a shared backbone, making policy adaptation harder to optimize, or under-utilize video information when guiding the action branch. In this work, we introduce DeVA, a Decoupled Video-Action model with specialized video and action experts, multi-level feature transfer, and physically salient guidance. DeVA transfers representations from multiple video layers to the action expert, enabling rich information exchange while making policy learning more tractable. It further supervises intermediate video features and the action stream with physically salient guidance (affordance/depth). Experiments on both simulation benchmarks and real-world deployment demonstrate strong performance with limited data, faster convergence than a unified architecture, and clear performance gains from physical guidance.

[CV-50] UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

链接: https://arxiv.org/abs/2607.24157
作者: Zhipeng Bao,Zhen Zhu,Nupur Kumari,Anurag Bagchi,Yu-Xiong Wang,Pavel Tokmakov,Martial Hebert
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single model produces diverse image-valued outputs through a unified multimodal interface. While diffusion-based systems dominate UVG due to strong quality and controllability, their iterative sampling incurs substantial inference latency, limiting practical deployment. To address these limitations, we propose UniGen-AR, a framework that pairs a general-purpose multi-modal language model (MLLM) with an efficient next-scale visual auto-regressive (VAR) decoder. This design retains the flexibility of MLLM-based conditioning while leveraging the sampling efficiency and latent unification properties of VAR models. In our framework, the MLLM encodes free-form instructions and control signals into a unified sequence, which guides the VAR decoder to generate image-valued outputs for over 15 tasks spanning four families. Empirically, UniGen-AR achieves up to 19 \times lower inference latency than diffusion-based baselines while maintaining or improving output quality. Our ablations further reveal that VQ-VAE tokenizer design, particularly codebook size and hierarchy, is a critical factor for VAR scalability in UVG. These results establish visual auto-regressive modeling as a compelling and efficient backbone for unified visual generation. Our project page is at this https URL.

[CV-51] Image Inpainting via Stochastic Dynamics

链接: https://arxiv.org/abs/2607.24140
作者: Jiaqi Kuang,Zihao Guo,Zhongmin Qian
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Image inpainting aims to recover missing regions while preserving structural consistency. We propose a non-parametric method without network training based on data-guided stochastic dynamics. Starting from a masked image, the missing pixels are evolved through a reverse-time stochastic differential equation with a kernel-weighted correction estimated directly from a reference dataset. This empirical correction guides the reconstruction toward high-density regions of the data distribution without training a neural network or fitting a parametric density model. Experiments on MNIST, Fashion-MNIST, and MVTec show that the proposed method outperforms Mean Fill, Telea, and Navier-Stokes inpainting in PSNR, SSIM, and visual quality. On CelebA, it remains competitive and produces plausible completions for structure-sensitive occlusions. These results demonstrate the effectiveness of empirical reference statistics as a non-parametric prior for image inpainting.

[CV-52] LoTA-N2N: Local Trace Adaptation for Zero-Shot Self-Supervised Image Denoising

链接: https://arxiv.org/abs/2607.24135
作者: Jintong Hu,Bin Xia,Junlin Liu,Jiayue Liu,Wenming Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22pages, 9 tables, 11 figures

点击查看摘要

Abstract:Single-image self-supervised denoising replaces unavailable clean targets with surrogate targets constructed from noisy observations. Its effectiveness therefore depends on how closely the surrogate objective remains aligned with supervised denoising, especially when noise is correlated, spatially nonstationary, or unknown. We express the discrepancy between a broad class of MSE-based self-supervised objectives and supervised MSE as a parameter-independent constant and a trace interaction between the surrogate-target residual and the prediction error. The corresponding gradient discrepancy is determined by the gradient of this interaction. This formulation provides a common view of paired-noise, blind-spot, weak-noise, re-corruption, and sub-image methods, while revealing that a small global interaction may conceal substantial positive and negative regional interactions through spatial cancellation. Building on these observations, we propose LoTA-N2N, a two-stage zero-shot adaptation framework. Stage 1 trains a denoiser on complementary sub-image pairs and freezes it to construct detached clean-sub-image proxies. Stage 2 estimates the residual–prediction interaction using these proxies and suppresses its patch-wise absolute magnitude. We show that the local construction prevents spatial cancellation and upper-bounds the magnitude of the corresponding global interaction. Experiments across natural, confocal, and X-ray images, complemented by iteration-matched controls, controlled noise shifts, and gradient diagnostics, show consistent gains over MSE-only adaptation under IID, spatially varying, and mixed noise. Overall, LoTA-N2N demonstrates that estimated local interaction and spatial cancellation control provide effective design principles for single-image self-supervised denoising without paired clean targets, repeated acquisitions, or a predefined re-corruption model.

[CV-53] ViDS: Video Diffusion Shader using 3D Face Tracking

链接: https://arxiv.org/abs/2607.24124
作者: Wenbo Ji,Davide Davoli,Zhe Chen,Liam Schoneveld,Matthias Nießner,Jiapeng Tang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages, 6 figures, and 8 tables. Project page: this https URL

点击查看摘要

Abstract:We introduce ViDS, a Video Diffusion Shader that leverages 3D face tracking for expressive and identity-preserving portrait animation. We first reconstruct the identity-specific 3DMM mesh from the reference image, and then animate it using expression and pose parameters from a driving video. Leveraging dense geometric cues from 3DMM normal maps, we employ a video diffusion model as a neural shader to synthesize lifelike portrait animations while preserving the appearance and identity of the reference image. We find that more accurate 3DMM tracking enables finer-grained expression control. We also introduce an autoregressive diffusion sampling process that extends generation beyond the model’s native window while reducing discontinuities between adjacent clips. Compared with prior diffusion-based approaches for portrait animation that rely on landmark-based conditioning or implicit motion latents, our method achieves more detailed and consistent expression and pose control while faithfully preserving identity and appearance. Detailed ablation studies validate the effectiveness of our design choices. Project page: this https URL

[CV-54] BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

链接: https://arxiv.org/abs/2607.24110
作者: Minchong Chen,Xiaoyun Yuan,Minyu Cao,Jianing Zhang,Jun Zhang,Shuyang Liu,Xiaokang Yang
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Optics (physics.optics)
备注: 13 pages

点击查看摘要

Abstract:Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment. In this paper, we propose BeyondFusion, a unified latent diffusion framework for calibration-free visible-guided infrared super-resolution and infrared-visible fusion tasks. The proposed framework supports both task-specific training and joint training where two tasks are optimized and executed as two readouts of the same generative process. Instead of relying on explicit registration or geometric warping, BeyondFusion introduces a cross-modal self-aligning (CMSA) module into the denoising U-Net. CMSA reorganizes infrared and visible latent tokens into a shared attention space to learn content-adaptive cross-modal correspondence during the denoising process. Together with misalignment augmentation module, the model is facilitated to exploit visible structural and semantic cues while preserving thermal consistency, enabling high-frequency infrared reconstruction and informative fused-image generation under uncalibrated conditions. Extensive experiments on public benchmarks and a mobile infrared-visible imaging system show strong performance across aligned inputs, low-resolution infrared observations, synthetic misalignments, and real mobile captures with unsynchronized sensors. Ablation studies, unified training analysis, and downstream pedestrian detection further validate the effectiveness of BeyondFusion for calibration-free multimodal imaging.

[CV-55] LU-500: A Logo Benchmark for Concept Unlearning

链接: https://arxiv.org/abs/2607.24101
作者: Keyu Li,Jin Gao,Jialing Zhang,Dequan Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Concept unlearning is increasingly used to limit the reproduction of protected or unsafe visual concepts in text-to-image models. Existing evaluations, however, mostly study targets that dominate the whole image, such as styles, broad object categories, or portrait-like identities, leaving company logos comparatively underexamined. Logos create a different failure mode: a small localized mark can carry the entire protected concept, must be visually precise to remain recognizable, and can be triggered implicitly by products, storefronts, packaging, or advertisements even when the word ``logo’’ is absent. We introduce LU-500, a logo-unlearning benchmark built from Fortune Global 500 companies to study this localized and semantically entangled setting. LU-500 contains nearly 10,000 curated text-query and logo-image pairs, with an explicit track (LUex-500) and an implicit contextual track (LUim-500). To avoid reducing the task to a binary detector score, we define a multi-grained protocol that evaluates both local logo removal and global image preservation in pixel and latent spaces. Experiments on representative inference-time methods, including NP, SLD, and SEGA, and compatible fine-tuning-based methods such as ESD and Forget-Me-Not, show that the evaluated methods struggle to remove logo evidence without changing non-target content. We further analyze ProLU, a prompt-space multi-agent baseline: it improves local erasure by removing logo-inducing semantics, but also illustrates why prompt filtering is not a substitute for weight-level disentanglement. Correlation analyses over logo area, location, and structural complexity suggest that future logo unlearning may need spatially aware controls, such as SSIM-guided constraints, rather than purely global concept suppression.

[CV-56] ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation

链接: https://arxiv.org/abs/2607.24098
作者: Yuanjia Li,Tianyang Xu,Tao Zhou,Zhangyong Tang,Xiao-Jun Wu,Josef Kittler
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Referring video object segmentation (RVOS) requires segmenting a target specified by natural language throughout a video. Recent agentic approaches combine multimodal large language models with promptable segmentation models to perform RVOS without task-specific training. However, most pipelines rely on one-shot spatial grounding followed by mask propagation, leaving both the initial prompts and temporal predictions largely unverified. We introduce ReflexTrack, a training-free, feedback-driven agent that closes this loop at both spatial and temporal levels. Mask-guided Spatial Refinement evaluates the mask induced by the current keyframe prompt and iteratively updates the bounding box together with positive and negative points, yielding a more reliable initialization. Video-level Mask Reflection assesses the complete mask sequence, localizes unreliable intervals, selects complementary repair keyframes, and generates candidate predictions through mask-guided re-propagation. Only candidates that provide a verified improvement are used to update the affected intervals, preserving reliable predictions elsewhere. All components remain frozen during inference. ReflexTrack achieves an overall \mathcalQ score of 69.7 on Ref-VPS and a \mathcalJ\mathcalF score of 67.2 on ReasonVOS. These results demonstrate that prediction-level feedback substantially improves the reliability of training-free RVOS.

[CV-57] Cascade Forgery Mining Network for Fingerprint Presentation Attack Detection

链接: https://arxiv.org/abs/2607.24090
作者: Hongyan Fei,Chuanwei Huang,Zheng Wang,Pengcheng Luo,Jingwei Li,Jufu Feng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Fingerprint Presentation Attack Detection (PAD) is a critical component of fingerprint identification systems, serving as a protective measure against unauthorized access. In this paper, we observe that different regions of a fingerprint image can exhibit varying Artifact Extraction Difficulty (AED), with high-AED regions requiring more sophisticated extraction mechanisms to capture more subtle discriminative evidence. To address this issue, we propose to quantify AED using local Gabor feature certainty and partition fingerprint images into multiple regions based on their respective AED values. We then propose an AED guided Cascade Forgery Mining Network (CFM-Net) that employs an adaptive-depth feature extraction architecture to detect more precise and comprehensive artifact evidence across regions with heterogeneous AED values. Furthermore, we introduce an Orientation Guided Adversarial Training (OGAT) module to filter out identity information from PAD features while preserving the integrity of original artifact evidence. Experimental evaluations on LivDet datasets demonstrate the superior performance of our approach compared to state-of-the-art methods and achieve significant improvement in the classification ability of high AED fingerprints.

[CV-58] When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

链接: https://arxiv.org/abs/2607.24077
作者: Marina Gardella(CB),Camilo Mari{ñ}o(UDELAR, CB),Diego Belzarena(UDELAR, CB),Ignacio Ram{í}rez(UDELAR),Gregory Randall(UDELAR),Jean-Michel Morel(LU - Hong Kong)
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
备注:

点击查看摘要

Abstract:Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on standard benchmarks. However, their suitability for archival transcription remains insufficiently understood. In this work, we benchmark traditional OCR systems and VLM-based approaches on the Berrutti dataset, a challenging collection of Uruguayan dictatorship-era documents derived from microfilm scans. While VLMs consistently outperform traditional methods in terms of Character Error Rate (CER) and Word Error Rate (WER), we show that these improvements hide a more complex picture. Through a detailed qualitative analysis, we uncover systematic failure modes that are invisible to standard metrics, including orthographic normalization, spurious content generation, and semantic substitutions that preserve fluency while altering meaning. Errors affecting named entities are particularly critical, as they can introduce substantial semantic distortions with minimal impact on CER and WER. These findings reveal a critical gap between quantitative OCR performance and transcription fidelity in real-world archival settings, and highlight the need for evaluation frameworks that go beyond character-level accuracy to capture the semantic reliability of generated transcriptions.

[CV-59] MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning ECCV

链接: https://arxiv.org/abs/2607.24064
作者: Tuan-An To,Yuk-Kwan Wong,Tuan-Anh Vu,Ziqiang Zheng,Sai-Kit Yeung
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted to The 19th European Conference on Computer Vision (ECCV) 2026

点击查看摘要

Abstract:Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality image-text pairs. However, the performance of VLMs often degrades in the video domain due to the essential need for temporal understanding and the scarcity of large-scale annotated video data. In this work, we focus on marine video understanding, which brings further challenges: first, it requires substantial domain expertise; and video VLMs usually struggle with localizing and interpreting critical information from marine videos, as the informative events are typically sparse, unpredictable, and unevenly distributed. To address these challenges, we carefully curate the first event-centric marine video understanding dataset called MarineEVT, which features 20K multi-task, video-level visual question-answering pairs spanning multiple dimensions of marine understanding and analysis. Meanwhile, based on MarineEVT, we decompose marine video understanding as an Event-centric Visual Tool-integrated Reasoning process EVT-R1 for short, where we leverage powerful visual tools to drive the model to localize and interpret critical information aligned with visual questions and human intent. To demonstrate its effectiveness, we compare EVT-R1 against 11 SOTA VLMs in different settings. EVT-R1 outperforms the top open-source and top commercial models by 5.22 and 11.09, respectively. MarineEVT and EVT-R1 lay the foundation for ecological discovery and marine education, fostering the development of VLMs capable of interpreting marine dynamics, reasoning about ecological interactions, and supporting sustainable ocean video understanding and analysis.

[CV-60] PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic Rectification

链接: https://arxiv.org/abs/2607.24052
作者: Xinxing Yu,Liying Yang,Hao Mo,Hui Ma,Fang Kai,Ajian Liu,Yanyan Liang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:High-curvature regions in 3D point clouds encapsulate critical fine-grained geometric semantics yet exhibit a distinct long-tail sparsity in their spatial distribution. The inherent limitations of polynomial volume growth in Euclidean space frequently render these intricate geometric features challenging to adequately resolve within a uniform-scale feature space. Consequently, these regions are frequently overshadowed by smooth global features dominated by low-curvature regions, thereby limiting the discriminative capacity of the network. To address this issue, we propose PointCHR, a curvature-aware hyperbolic rectification (CHR) for point cloud analysis. Utilising the property of exponential volume expansion in the vicinity of hyperbolic manifolds, CHR presents a learnable curvature-guided radial rectification mechanism. By adaptively projecting high-curvature points towards boundary regions endowed with larger effective embedding capacities, PointCHR effectively mitigates the representation crowding problem inherent in Euclidean settings. Extensive experimentation has demonstrated that PointCHR significantly enhances the ability of backbone to capture fine-grained geometric details, achieving state-of-the-art performance across multiple benchmarks.

[CV-61] Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

链接: https://arxiv.org/abs/2607.24027
作者: Haopeng Li,Yitong Li,Junsong Chen,Tian Ye,Haozhe Liu,Jincheng Yu,Duomin Wang,Ruihua Zhang,Zeke Xie,Enze Xie,Song Han
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: technical report

点击查看摘要

Abstract:Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing methods struggle to sparsify attention both efficiently and accurately for two reasons: (1) Rigid, unpredictable, and costly routing: selecting a fixed fraction of top-ranked blocks by proxy score imposes fixed budgets, whereas retaining blocks to reach a target cumulative proxy probability mass yields dynamic but potentially imbalanced budgets; both incur non-negligible overhead from computing and materializing proxy scores. (2) Lossy keep-or-drop sparsification: unselected blocks are discarded entirely, degrading accuracy under aggressive sparsity. These limitations motivate cheaper dynamic-budget routing while limiting accuracy degradation. In this paper, we introduce training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention. The core of Sol-Attn is on-the-fly block thresholding with proxy-score reuse, which selects critical blocks by comparing block proxy scores against a threshold during online softmax. This design enables dynamic yet controllable block budgets without materializing the proxy map, while directly reusing the proxy scores of unselected blocks to approximate their contribution. Experiments across image and video generation tasks show that Sol-Attn advances the quality-efficiency frontier of training-free sparse attention, delivering 2.1 times and 2.3 times end-to-end speedups for video generation and editing, respectively, while preserving visual quality.

[CV-62] A Unified Stereo Geometry Estimation Framework for Disparity and Surface Normal

链接: https://arxiv.org/abs/2607.24024
作者: Qizhe Wei,Xianda Guo,Shaocong Xu,Hong Li,Runyi Yang,Hao Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Stereo matching and surface normal estimation are fundamental tasks in 3D vision. However, existing feed-forward stereo methods still struggle to produce reliable predictions in challenging regions, mainly due to the lack of strong geometric priors. In this paper, we propose \textbfGeoStereo , a unified stereo geometry estimation framework that leverages powerful diffusion priors to jointly predict disparity and surface normals. Specifically, GeoStereo couples a feed-forward stereo matching pipeline with a diffusion-based normal estimation branch. To enable effective interaction between the two tasks, we introduce a disparity to normal initialization strategy and construct a warp to left-view condition for the diffusion process. This coupled design allows the diffusion branch to provide strong structural priors that enhance disparity estimation in ill-posed regions, while the feed-forward branch offers reliable geometric guidance for accurate normal prediction. Extensive experiments show that GeoStereo performs reliably in challenging scenarios, including low-light environments, highly reflective surfaces, and transparent objects. Under zero-shot settings, it achieves Rank-1 disparity estimation on multiple benchmarks, including KITTI and NYUv2, and delivers the best normal estimation accuracy on many real indoor benchmarks, such as iBims-1 and ScanNet. Project page: this https URL

[CV-63] Disentangling Semantic Attention from Structural Bias in the Attention Manifold

链接: https://arxiv.org/abs/2607.24017
作者: Pengkun Jiao,Bin Zhu,Jingjing Chen,Yu-gang Jiang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed “register” or “Visual Attention Sinks.” While existing inference intervention methods attempt to identify these sink tokens and redistribute their attention weights, such approaches typically treat these tokens in isolation and suffer from computational inefficiency. Instead, we reframe this phenomenon as a generalized textual bias exerted over visual features that extends beyond isolated sink tokens. From this perspective, a pervasive structural bias leads to the dilution of the semantic visual signal, precipitating multimodal hallucinations as the model prioritizes linguistic priors over valid visual evidence. To address this limitation, we introduce Saliency-guided Purification and Adaptive Redistribution (SPAR), a training-free, plug-and-play intervention. SPAR mitigates this generalized textual bias by purifying structural noise and subsequently redistributing the reclaimed attention budget to the most informative visual regions. Comprehensive evaluations across a diverse spectrum of hallucination benchmarks demonstrate that SPAR effectively restores authentic visual grounding with negligible computational overhead.

[CV-64] DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models

链接: https://arxiv.org/abs/2607.24016
作者: Xin Jiang,Hao Tang,Junyao Gao,Meiqi Cao,Fei Shen,Dongming Zhang,Yongdong Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent advances in generative models have shifted AI-generated image detection from identifying easily distinguishable, fully synthetic images to identifying highly realistic content generated by both modern generation and manipulation pipelines. However, existing detection benchmarks are often built with outdated generative models and primarily emphasize full-image synthesis, creating a growing mismatch between benchmark data and the images encountered in real-world generation and editing scenarios. To bridge this gap, we introduce DailyBench, a high-quality unified benchmark for evaluating whether AI-generated image detectors can generalize across both modern full-image synthesis and object-level manipulation. DailyBench contains two complementary subsets: FakeBench, which includes high-quality images synthesized by recent open-source and commercial generative models, and ManipulationBench, which introduces challenging object-level edits applied to real images using advanced image-conditional models. This design makes DailyBench a realistic testbed for studying both generator-level generalization and manipulation-aware detection under subtle local edits. Experiments on DailyBench reveal substantial robustness gaps in current detectors: methods reporting 91-96% balanced accuracy on GenImage drop to 60-76% on FakeBench and 54-66% on ManipulationBench. These results show that existing detectors remain poorly generalized to realistic synthesis and manipulation, highlighting DailyBench as a rigorous testbed for developing robust and manipulation-aware AI-generated image detection methods. The project is available at this https URL

[CV-65] AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

链接: https://arxiv.org/abs/2607.24013
作者: Hengyuan Zhang,Jingna Sun,Meiguang Jin,Junfeng Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 5 figures

点击查看摘要

Abstract:Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often compromise quality through restrictive architectural choices, such as causal attention and short temporal horizons, or by reducing model capacity and resolution. Without such compromises, we propose AptAvatar, a 14B-parameter long-form audio-driven avatar generation framework that delivers fast and expressive inference. For efficiency in production-level applications, AptAvatar addresses the extreme two-step generation challenge. To bridge the gap between the multi-step teacher model and the two-step student model, we introduce Endpoint-Anchored Distribution Distillation. It augments vanilla distribution matching with a dedicated Anchor Score Estimator trained on the trajectory-endpoint distribution defined from a frozen pretrained 4-step bridge generator. This provides an attainable endpoint-level anchor for the evolving two-step student. To improve long-horizon consistency, we further introduce Self-Generated History Replay, which reuses cached outputs from earlier generator checkpoints as history conditions during chunk-wise training. This approximates inference-time conditioning on self-generated histories without costly online rollouts, mitigating quality degradation from accumulated history errors. Extensive experiments demonstrate that AptAvatar generates vivid 720p long-form avatar videos with only 2 NFEs, achieving a 60x speedup while preserving visual fidelity and long-horizon identity. Code is available at this https URL

[CV-66] Structural Loss Metrics for Tensor Approximation via Matrix Low-Rank Approximation

链接: https://arxiv.org/abs/2607.24009
作者: Hiroki Hasegawa
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 4 figures

点击查看摘要

Abstract:Matricized low-rank approximation via SVD is a standard surrogate for tensor decompositions, but entry-wise reconstruction error fails to capture multiway geometric degradation. Under an orthogonal Tucker model, we characterize this degradation using two metrics: cross-mode Direction Loss, measuring geometric subspace deviation from rank truncation and noise rotation, and Interaction Loss, quantifying multilinear interaction distortion in the core tensor. We prove that squared relative reconstruction error orthogonally decomposes into interaction loss and out-of-subspace energy, and derive a Wedin-type bound establishing the stability of a plug-in Direction Loss estimator. Experiments on synthetic and hyperspectral datasets demonstrate that nearly identical reconstruction errors can yield markedly different structural-loss profiles; hyperspectral patches with comparable reconstruction errors exhibit up to a 4.6-fold difference in Direction Loss, correlating with severe visual blurring.

[CV-67] Low-light Image Enhancement via Multi-scale Attention combined with Fourier Transform

链接: https://arxiv.org/abs/2607.24002
作者: Wenbin Du,Jian Long,Zhu Cao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 27 pages, 8 figures

点击查看摘要

Abstract:Low-light image enhancement (LLIE) aims to improve image quality and clarity in diverse and demanding low-illumination environments. However, existing deep learning-based LLIE methods struggle to accurately capture real-world illumination and restore texture details, largely because their algorithmic strengths remain underutilized. To address these issues, we present a supervised frequency domain deep learning network for LLIE, named multi-scale attention combined with the Fourier transform (MSFT) which adopts a U-shaped, one-stage architecture that infuses guidance from low-light images into the network by channeling it through multi-scale attention. We further fuse the amplitude information from priori channels with that of the low-light image in MSFT’s self-created module, and carry out multi-scale guidance along with the network. Subsequently, to better enhance the faint feature, such as fine content and textures, and to better fuse global context confidence in the decoding stage, we separately introduce a multi-shape synergistic attention and a lightweight network that effectively integrate information in high-dimensional space to embed into the superlative feature space channel containing rich texture information. Extensive experiments conducted on LOL, SID, SMID, and SDSD datasets demonstrate that MSFT significantly outperforms state-of-the-art competitors. For example, compared with Retinexformer, our method achieves a peak signal-to-noise ratio of up to 41.76 decibels on the SDSD-outdoor dataset with an increase of 11.92 decibels and a structural similarity index of 0.988 with a 13.80% improvement.

[CV-68] Effective Receptive Field Ordering Matters for Infrared Small Target Detection

链接: https://arxiv.org/abs/2607.23994
作者: Guoyi Zhang,Yanjin Du,Zhengyao Zhao,Tongsu Zhang,Guangsheng Xu,Siyang Chen,Xiangpeng Xu,Han Wang,Xiaohu Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:In this work, we investigate a previously unexplored architectural dimension for infrared small target detection: the organization of effective receptive fields (ERFs) during feature refinement. Unlike existing approaches that primarily improve individual feature operators, we argue that ERF organization constitutes an architectural dimension independent of receptive field design itself, and formulate deep feature transformation as a progressive residual correction process, from which a theoretical framework for ERF scheduling is established. Specifically, we reveal that ERF refinement is governed by two fundamental properties: scale-frequency correspondence, which aligns different ERF scales with distinct residual frequency characteristics, and nonlinear non-commutativity, which makes different ERF orderings produce fundamentally different refinement trajectories. Together, these properties show that ERF organization, rather than ERF scale alone, governs refinement dynamics. Guided by these principles, we propose Receptive Field Ordering Network (RFONet), which realizes hierarchical ERF scheduling through a multigrid-inspired V-cycle strategy using only standard 3\times3 convolutions. RFONet achieves state-of-the-art performance on multiple benchmarks with only 1.16M parameters and over 157 FPS inference speed. Beyond empirical performance, our theoretical analysis provides theoretical guarantees for stable residual refinement under perturbations, frequency shifts, and partial occlusions, which are consistently reflected in superior noise robustness and cross-dataset generalization. Finally, our framework reformulates ERF organization as a task-dependent optimization objective, providing a principled foundation for future adaptive receptive field scheduling.

[CV-69] Multimodal Semantic-Probabilistic Objectness for Open World Object Detection

链接: https://arxiv.org/abs/2607.23981
作者: Weijun Tian,Rui Liu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 16 pages, 3 figures, 4 tables

点击查看摘要

Abstract:Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabilistic objectness in the decoder-query space. However, visual objectness alone cannot determine whether an object-like query corresponds to a hard known instance, an unseen-category object, or background clutter, resulting in an ambiguous known-unknown decision boundary. We propose MSPO, a lightweight semantic calibration framework that augments PROB with task-aware known-category language priors while preserving its detector architecture and incremental learning protocol. For each currently known category, MSPO constructs an extended text description covering category attributes, visual appearance, typical scenes, and functional usage, and encodes it using a frozen CLIP text encoder. Decoder query features are projected into the same semantic space to estimate their support from the current known-category semantics. This semantic evidence is fused with PROB’s visual objectness to calibrate known and unknown predictions without turning OWOD into open-vocabulary classification. Importantly, MSPO never uses future-category names, and all unseen categories remain unnamed during evaluation. Experiments on M-OWODB and S-OWODB show that MSPO improves the strong PROB baseline on the main aggregate metrics while retaining competitive unknown recall. It also improves early unknown-confusion metrics and raises PASCAL VOC final mAP by up to 2.7 points. These results demonstrate that known-category language semantics provide an effective calibration signal for probabilistic objectness under the standard OWOD setting.

[CV-70] Mutual Modality Trust with Lightweight Reconstruction Regularization for Fine-grained Tire Pattern Recognition

链接: https://arxiv.org/abs/2607.23979
作者: Jianning Yang,Jie Fang,Xinda Ma,Zirui Song,Dianwei Wang,Nan Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Visual tire recognition serves as a core supporting technique for vehicle safety monitoring, autonomous driving perception and automated automotive maintenance. Existing fine-grained tire recognition techniques suffer from three prominent limitations. They tend to depend on only one visual source, lack the capacity to jointly model spatial and frequency cues for minute tread texture extraction, and suffer severe overfitting given limited annotated tire imagery. This paper proposes a lightweight fine-grained tire pattern recognition method incorporating dual-branch independent inference and enhanced feature fusion to boost recognition performance. The framework employs two task-specialized branches dedicated to tire surface and tread indentation, respectively, to extract modality-specific discriminative features. Each branch conducts independent prediction, while cross-branch feature fusion exploits Mutual Modality Trust (M ^2 T) to realize complementary feature enhancement across two modalities. Besides, a frequency-domain hierarchical guidance module is devised, which leverages bandpass filters to decompose feature maps into high- and low-frequency components and enables fine-grained cross-layer feature modulation. Furthermore, a Lightweight Reconstruction Regularization (LR ^2 ) is introduced to retain abundant intrinsic information within feature embeddings, substantially improving feature stability and recognition robustness under limited labeled training data. In addition, we establish a surface-indentation multi-source dataset namely MTire299 for fine-grained tire tread recognition, which covers 299 categories with a total of 14795 paired image samples. Extensive experiments conducted on two public tire datasets validate the superiority and efficacy of the proposed algorithm.

[CV-71] Color Fundus Photography Analysis: Co-evolution of Data Preprocessing and Modeling toward Multimodal AI

链接: https://arxiv.org/abs/2607.23972
作者: Yu Li,Wengan He,Wenhui Xu,Lihong Jiang,Fan Xiao,Zhuohang Huang,Yuanzhu Liang,Jiayi Liu,Yuxi Chen,Yongsheng Luo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Survey paper, 77 pages, 18 figures, 2 tables

点击查看摘要

Abstract:Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence. This review provides an integrated overview of CFP AI through the interplay of dataset evolution, preprocessing paradigms, and modeling frameworks. We show that CFP datasets have evolved from small single-center collections with task-specific labels to large multi-center resources featuring multimodal pairings and longitudinal clinical records. Preprocessing has progressed from conventional image enhancement to neural data-engineering pipelines, hardware-aware token optimization, and self-supervised imputation for incomplete electronic health records (EHRs). Meanwhile, modeling has advanced from convolutional neural networks (CNNs) to vision foundation models, state space models (SSMs), and multimodal expert architectures. At the multimodal frontier, CFP is increasingly integrated with EHRs and longitudinal patient information, enabling more comprehensive clinical reasoning beyond isolated image analysis. We conclude that future progress depends on the collaborative optimization of datasets, preprocessing, and multimodal modeling, providing a roadmap toward robust clinical deployment, improved cross-domain generalization, and resource-efficient edge intelligence.

[CV-72] Development of Vision-Language Model-based GNSS Spoofing Detection for Autonomous Vehicle Navigation

链接: https://arxiv.org/abs/2607.23962
作者: Mohammed Aldeen,Muhammad Sami Irfan,Sagar Dasgupta,Long Cheng,Mizanur Rahman,Mashrur Chowdhury
类目: Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
备注:

点击查看摘要

Abstract:Autonomous vehicles (AVs) depend on Global Navigation Satellite Systems (GNSS) for localization and navigation, making them vulnerable to spoofing attacks that can covertly redirect vehicles or induce unsafe maneuvers. In this paper, we develop the first Vision-Language Model (VLM)-based framework for GNSS spoofing detection for autonomous vehicles by fusing front-camera visual data with in-vehicle sensor readings (e.g., speed, acceleration, yaw rate) against GNSS-derived maneuvers. Our approach introduces a three-stage fine-tuning process that first grounds visual cues, and then calibrates sensor data within a shared semantic space to detect discrepancies between predicted and GNSS-derived maneuvers across three attack scenarios. We also generated an independent real-world dataset by driving an instrumented vehicle on public roads in Tuscaloosa, Alabama, equipped with time-synchronized GNSS, IMU, and camera logs to validate cross-regional generalization of our fine-tuned model on unseen data from training data. On this dataset, we then generated intelligent spoofing attacks, including trajectory mirroring with road-network snapping for wrong-turn attacks, position freezing for overshoot scenarios, and drift generation for stop attacks. On this validation dataset, the zero-shot VLMs baseline F1-score ranges from 23% to 32%, whereas our fine-tuned model achieves an F1-score ranging from 94% to 95%. Results show that our VLM-based approach correctly classified every wrong-turn and stop attacks, and attains 88%-93% accuracy for overshoot attacks. Furthermore, we introduce an adaptive inference policy that reduces VLM invocations to 14% (~86% computational reduction) and yields 65ms-73ms per 4s window. These results point to a practical, on-road layer of defense that complements signal-level integrity checks with the use of VLMs.

[CV-73] RODR: Riemannian Orthogonally Decoupled Regularization for Disentangled Manifold Representation

链接: https://arxiv.org/abs/2607.23958
作者: Jiayu Zhu,Wenlai Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 4 figures

点击查看摘要

Abstract:Point cloud denoising is essentially a geometric recovery task that aims to reconstruct the intrinsic structure of a smooth 2D Riemannian manifold embedded in R^3 from noisy, discrete ambient-space samples. Despite the remarkable progress of modern manifold-aware encoders and generative transport models in geometric representation learning, a fundamental objective-geometry mismatch remains underexplored. Theoretically, we identified that this mismatched coupling leads to geometric gradient interference, where conflicting optimization objectives result in structural degradation and point clustering. We introduce Riemannian Orthogonally Decoupled Regularization (RODR) to reformulate the optimization trajectory by disentangling the normal (fitting) and tangential (distribution) components. Guided by a vector-attention and entropy-aware adaptive strategy, RODR effectively preserves high-fidelity geometric details while maintaining sampling uniformity. Experiments demonstrate that RODR reaches performance comparable to state-of-the-art baselines and suggests improved distribution regularity and reduced local aggregation effectively. Our work establishes a generic and interpretable framework for disentangled geometric optimization in point cloud processing.

[CV-74] mePLE: Rethinking Temporal Representation for Video Temporal Grounding

链接: https://arxiv.org/abs/2607.23951
作者: Yuhui Zeng,Xinyu Mao,Xiaokun Liu,Xin Tao,Jinfa Huang,Jiayi Ji,Xiawu Zheng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 25 pages, 13 figures

点击查看摘要

Abstract:Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this interval indirectly through two endpoint outputs, represented either as discrete timestamp tokens or continuous boundary coordinates. These formulations differ in how endpoints are encoded, but not in what is predicted: the event interval remains a derived object, while interval validity, duration, and interval-level similarity are handled only implicitly. We propose TimePLE, which reformulates VTG from endpoint prediction to interval-native grounding by predicting a single joint distribution over valid temporal intervals. TimePLE maps each interval to a point in a canonical position-duration square, where every support point corresponds to a valid span and neighboring points represent geometrically similar intervals. Given a video and query, the VLM generates a single latent |TIMESPAN| token whose hidden state is decoded into a joint interval distribution, refined through duration-aware coordinate correction, and converted into continuous boundaries. The same interval representation is used to encode input temporal anchors, aligning video-side temporal evidence with output-side span prediction. To reliably align the latent span representation with complete event intervals, we curate 90K-scale grounded samples and human-verify 3K-scale benchmark annotations. Experiments across four VTG benchmarks show that TimePLE consistently outperforms endpoint prediction baselines, achieving an average mIoU of 58.9, with clear gains on short-duration and medium-duration events.

[CV-75] Manifold-Constrained Noise Optimization for Diverse Diffusion Sampling

链接: https://arxiv.org/abs/2607.23937
作者: Qitan Shi,Cheng Jin,Ziyuan Liu,Yuantao Gu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 14 figures, 4 tables

点击查看摘要

Abstract:Few-step distilled diffusion models generate high-quality images quickly, but often lose per-prompt diversity, producing near-identical samples across random seeds. Optimizing the initial noise at inference time offers an appealing way to recover this diversity, yet existing methods directly update the initial noise in an unconstrained Euclidean space, ignoring both the geometry of the Gaussian prior and the model’s sensitivity to noise frequencies. They therefore introduce auxiliary quality-control objectives to maintain generation fidelity, adding compute and weighting hyperparameters while still requiring conservative updates to prevent degradation. In this work, we propose MoNO, a training-free method that performs Manifold-constrained Noise Optimization on a low-dimensional, quality-stabilizing noise manifold. MoNO sequentially optimizes each new initial noise so that its predicted visual feature complements previous generations, while Riemannian updates on an affine low-frequency sphere preserve prior likelihood and fix unstable high-frequency components by construction. This enables large geodesic steps, removes the need for auxiliary quality-control objectives, and converges in far fewer iterations than prior noise-optimization methods. Experiments with multiple distilled text-to-image diffusion models show that MoNO consistently improves per-prompt diversity while maintaining image quality.

[CV-76] DuoAD: Leverag ing [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection

链接: https://arxiv.org/abs/2607.23924
作者: Jyun-Ze Tang,Po-Han Huang,Ming-Ching Chang,Chih-Fan Hsu,Jeng-Lin Li
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Code: this https URL

点击查看摘要

Abstract:Vision foundation models have enabled strong training-free anomaly detection (AD). However, most existing approaches rely primarily on independent local patch features, leaving the global contextual information encoded by Vision Transformers (ViTs) underexploited. In this work, we identify the dual characteristics of the ViT [CLS] token: its embedding provides anomaly-invariant global semantic representation, while its attention maps implicitly highlight spatially abnormal regions. Building on this observation, we propose a fully automated AD framework leveraging global context to remove manual tunings. Our framework introduces (1) an automatic augmentation selection strategy driven by [CLS]-level semantic consistency, and (2) an attention-guided feature reweighting mechanism that dynamically adjusts patch contributions according to [CLS] attention saliency. By integrating these components over multi-level features, our method achieves stable anomaly scoring and precise localization without training or parameter tuning. Under the one-shot setting, it achieves Image-AUC scores of 97.7%, 93.2%, and 84.5% on MVTec-AD, VisA, and Real-IAD. Using a single fixed configuration across categories, backbones, and datasets, the method establishes a new state-of-the-art for plug-and-play, training-free anomaly detection while maintaining strong robustness and practical scalability.

[CV-77] DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering

链接: https://arxiv.org/abs/2607.23921
作者: Yue Zhang,Xiangyu Li,Wanshu Fan,Xin Yang,Dongsheng Zhou
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by CGI

点击查看摘要

Abstract:Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered significant attention due to its practical applications. In this paper, we introduce the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. Specifically, we propose a question-guided dynamic regional-level module (QGDR) that combines complementary image context through ROI Align and text content, enabling precise localization of text-related visual content. Moreover, we present a cross-modal multi-scale aggregation module (CMA) that enhances feature fusion between pixel-level and region-level features, facilitating the effective localization of visual content associated with grounded answers. Furthermore, we fuse the located visual content with text features to locate the region and provide answers to questions posed about the image. Experimental results demonstrate that our DDVT outperforms state-of-the-art methods on several widely-used benchmarks.

[CV-78] What Can I Edit? Open-Ended Strategy Discovery and the Emotion Editability Landscape

链接: https://arxiv.org/abs/2607.23920
作者: Qing Li,Zeyu Dong,Yin Cui,Chuan Yan,Xiaojiang Peng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Emotional image editing requires more than applying affective filters or modifying predefined visual factors: an effective edit must identify what a particular image can afford for a target emotion. Existing affective image manipulation methods, including recent agentic variants, largely operate within bounded strategy spaces based on predefined factor taxonomies, knowledge libraries, or conventional editing templates, and therefore often miss image-specific, context-grounded strategies. We introduce EmoScope, a multi-agent framework that reframes the task from “how should I edit?” to “what can I edit?” EmoScope first discovers an image-specific editable space through emotion-conditioned affordance reasoning, then uses a semantic hierarchy of anchors, variables, and context to balance content consistency and emotional expressiveness before executing and verifying the edit. Because its plans are expressed as image-specific affordances rather than retrieved templates, EmoScope also exposes the editing strategy as an interactive surface for user refinement at the plan level. In a large-scale human evaluation covering all eight Mikels emotion categories, with 4,693 valid responses across 1,824 pairwise questions, participants preferred EmoScope over two competitive baselines by 88.1% on average. Attribution analysis further shows that EmoScope selects target-emotion-adaptive strategies rather than applying a uniform template. The same affordance-level plan also supports lightweight user refinement in an interactive pilot. Finally, we show that classifier-based metrics exhibit emotion-conditional blind spots toward non-stereotypical, context-grounded edits, and present a relative content-emotion preference-affinity landscape showing that EmoScope’s advantage varies systematically across image-emotion combinations.

[CV-79] Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention ECCV

链接: https://arxiv.org/abs/2607.23917
作者: Sounak Mondal,Dimitris Samaras,Gregory Zelinsky,Minh Hoai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: To appear in European Conference on Computer Vision (ECCV) 2026

点击查看摘要

Abstract:We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as a discriminative task over predefined categories, we formulate it as a generative learning problem: training a model to produce free-form descriptions that capture the rich nuances and open-ended nature of human intentions beyond fixed labels. To this end, we introduce Gazette, the first gaze-to-text decoding framework. Based on multimodal large language models (MLLMs), Gazette learns to decode gaze scanpaths into natural language for goals that may extend beyond categorical labels and require articulation in natural language. To help Gazette filter out individual differences in gaze behavior and learn the goal-specific spatiotemporal dynamics crucial for generating accurate natural language goal descriptions, we propose a novel strategy that leverages the encyclopedic knowledge and reasoning abilities of a large language model to synthesize natural language explanations of goal-directed attentional behavior called think-aloud transcripts. Instruction tuning on these synthetic narratives allows Gazette to achieve state-of-the-art performance in gaze decoding across multiple tasks, demonstrating its generalizability and versatility, thereby enabling gaze to serve as a powerful, non-intrusive cue for inferring human goals and intentions in diverse scenarios.

[CV-80] Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets

链接: https://arxiv.org/abs/2607.23908
作者: Syed Roshaan Ali Shah,Kristof Van Tricht,Christina Butsko,Jeroen Degerickx,Zoltan Szantoi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:High quality reference data remain a critical bottleneck for crop-type mapping at any spatial and temporal scale. Operational systems such as WorldCereal aggregate labels from heterogeneous sources such as parcel registers, national databases, field surveys, and map-derived products, each with their own biases, coverage gaps and unknown label noise. Simple global rules are inadequate, since crop phenology and observation conditions vary strongly across regions and seasons. In this study, we focus on a single, operationally relevant question: whether embeddings produced through geospatial foundation models are a viable basis for cleaning the reference data. We propose a practical, locality-aware, embedding-based anomaly (EBA) detection framework that operates on the embeddings of a pretrained Earth-observation encoder. We score each labelled sample against other samples of the same crop in the same area using a pretrained embedding, flag the ones that stand out, and test whether removing or down-weighting them before training yields a better model. We establish that the flagged points are genuinely mislabelled or misplaced in two independent ways: against synthetic ground truth, the detector concentrates injected label errors 2.5-5x above chance in its flagged set (detection AUROC up to 0.84); and on real data, a model-independent test shows that removing or confidence-weighting the flagged held-out points raises measured accuracy in trained models, for both crop type and land cover. Acting on the flags then improves the WorldCereal crop-type model across five macro-regions, evaluated on a fixed held-out split under three views. We find conservative cleaning helps while over-cleaning hurts. The EBA detector approach is designed to be reproducible and extensible, and can serve as a template for cleaning large, noisy Earth observation reference datasets beyond crop mapping. Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.23908 [cs.CV] (or arXiv:2607.23908v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2607.23908 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Syed Roshaan Ali Shah Mr [view email] [v1] Mon, 27 Jul 2026 00:52:52 UTC (2,152 KB)

[CV-81] Long-Tailed Medical Image Classification

链接: https://arxiv.org/abs/2607.23883
作者: Nathanael Ren,Saagar Arya
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:In this paper, we examine the difficulties of using standard techniques for medical image classification due to long-tailed distributions (wherein rarer conditions have very few samples) resulting in bias towards diagnosing common diseases and away from rarer diseases. We then discuss and implement deep learning models with techniques such as augmentation to minimize error, especially from rarer diseases. We evaluate various different models with AP, F1 score, AUROC, and loss (all on the validation set). We conclude with the promising results from our best model, and potential applications in the healthcare space.

[CV-82] Head Avatars with Dynamic Explicit Hair

链接: https://arxiv.org/abs/2607.23861
作者: Vanessa Sklyarova,Haonan Chen,Berna Kabadayi,Tobias Kirschstein,Zicong Fan,Xi Wang,Gerard Pons-Moll,Matthias Nießner,Marc Pollefeys,Michael J. Black,Justus Thies
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: Project page: this https URL

点击查看摘要

Abstract:We present DynHair, a novel method for tracking and modeling dynamic hair for human head avatars. From video input, we reconstruct a dynamic head avatar with an explicit strand-based hair representation using structured 3D Gaussian Splatting. In contrast to the face region of human head avatars, which can be modeled with 3D Gaussians that are attached or generated with respect to some expressive 3D head model, hair is particularly challenging as it exhibits dynamic motion effects. Therefore, we present a novel method that models the dynamic deformations of the hair strands using a temporal network that is conditioned on angular velocity and acceleration of the head, as well as relative gravity. Specifically, an LSTM encodes the motion history and modulates per-point strand features via FiLM conditioning which further used by MLP to produce physically plausible displacements to canonical hairstyle. We jointly optimize this motion and appearance representation of the hair, with a 3DGS-based representation of the face-region, via differentiable Gaussian splatting with photometric, geometric, and physics-based supervision. As a result of our method, we retrieve hair tracking of the training video data and an animatable head avatar with controllable hair dynamics. In our experiments, we demonstrate state-of-the-art performance in terms of hair dynamics, temporal consistency, and generalization across subjects.

[CV-83] Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation

链接: https://arxiv.org/abs/2607.23860
作者: Mihai Suteu,Ovidiu Serban
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deep ensembles provide the most reliable uncertainty estimates in deep learning, but their cost grows linearly with the number of members. Implicit ensembles lower this cost by sharing a single backbone across members. Member diversity is a primary determinant of ensemble quality, yet no implicit ensemble can shape it during training; existing methods fix it at initialisation or build it into the architecture. We introduce \sigma N-Ens, a normalisation-based implicit ensemble that treats each member as a task in a multi-task architecture and modulates the shared backbone through sigmoid-bounded scalers. We also introduce a softmax-temperature regulariser, which shapes the equilibrium level of sharing between members and traces the accuracy-calibration frontier. Because only normalisation layers are replicated, the mechanism can wrap convolutional and transformer backbones alike, also allowing pretrained models to be adapted through a short fine-tune. We frame the epistemic uncertainty such an ensemble expresses as modulation uncertainty, and explain why its calibration holds under input corruption, and why its out-of-distribution detection is weaker. Our method is evaluated across ResNets and transformers on CIFAR-10/100, ImageNet and SST-2. \sigma N-Ens matches or outperforms deep ensembles at a fraction of their parameter cost, scales with ensemble size where partitioning methods collapse, and maintains calibration under distribution shift.

[CV-84] OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

链接: https://arxiv.org/abs/2607.23855
作者: Jun Zhan,Chen Yang,Yitian Gong,Donghua Yu,Kuangwei Chen,Wenbo Zhang,Kexin Huang,Qi Luo,Zhe Xu,Ying Zhu,Jin Wang,Tengyue Zhang,Qi Chen,Cheng Chang,Songlin Wang,Junqi Dai,Jiasheng Ye,Xiaogui Yang,Tianyi Liang,Xiangyu Peng,Zhaoye Fei,Shimin Li,Qinyuan Cheng,Xie Chen,Xinchi Chen,Xipeng Qiu
类目: ound (cs.SD); Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 2 figures, 6 tables

点击查看摘要

Abstract:Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1

[CV-85] OmniCache: Multidimensional Hierarchical Feature Caching For Diffusion Models

链接: https://arxiv.org/abs/2607.23844
作者: Zhaoyuan He,Muhammad Muaz,Lili Qiu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:High-resolution image and video diffusion models, including SD3, FLUX, and recent video diffusion transformers, have substantially improved generative quality but remain expensive at inference time because they repeatedly evaluate attention-heavy denoisers over many sampling steps. We address this inefficiency by exploiting redundancy in intermediate diffusion features rather than changing model weights or retraining. We identify four complementary redundancy sources in image and video generation: intra-frame, inter-frame, motion, and denoising-step redundancy. Based on this analysis, we propose OmniCache, a unified hierarchical caching framework that performs multidimensional feature reuse through Token Cache, Frame Cache, Block Cache, and Layered Cache. Unlike token-merging baselines that average matched features, OmniCache uses similarity matching to select cacheable features, skips redundant computation, and restores positionally consistent cached activations, preserving feature order and spatial-temporal structure. The resulting framework reuses spatial features in temporal layers and temporal features in spatial layers, while Layered Cache captures cross-step redundancy at the model-layer level. Across SD3, SVD-XT, and Latte, OmniCache reduces inference latency by up to 35%, 25%, and 28%, respectively, while maintaining visual fidelity and motion coherence in a training-free setting.

[CV-86] STEER: Steerable Dyadic Head Avatars

链接: https://arxiv.org/abs/2607.23840
作者: Kartik Teotia,Helge Rhodin,Hyeongwoo Kim,Marc Habermann,Christian Theobalt
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL

点击查看摘要

Abstract:Facial movement and expression are central to face-to-face communication, conveying turn-taking, attention, agreement, and engagement alongside speech. While speech-driven facial animation has made strong progress in lip synchronization and audio-conditioned motion generation, most methods treat conversational behavior as an emergent byproduct of audio, or expose only coarse sequence-level affect control. As a result, key non-verbal channels such as gaze contact and aversion, rhythmic head motion, and emotion remain difficult to explicitly control. We present STEER, a controllable 3D dyadic motion prior for reactive conversational head avatars. STEER factorizes conversational behavior into explicit controls for gaze, head rhythm, and emotion, allowing users to steer how an avatar listens, reacts, and engages with a conversation partner. Since temporally aligned annotations for these behaviors are not available in public dyadic corpora, we introduce a tracking and annotation pipeline that recovers behavioral pseudo-labels from in-the-wild dyadic video. A causal flow-matching transformer then learns partner-aware target motion conditioned on audio, partner motion, emotion and the proposed behavioral controls. We further embed STEER in a photorealistic avatar pipeline by extending a Universal Gaussian Head-Avatar Prior with a learned mapping from tracked parametric motion into its avatar-driving space. This enables controllable animation of high-fidelity Gaussian head avatars without re-training the underlying avatar model. STEER outperforms recent dyadic motion baselines on motion quality, dynamics, and diversity, remains competitive on partner coupling, and enables gaze, head-rhythm, and emotion edits together with an interactive live deployment. We make our code and dataset annotations available at our webpage.

[CV-87] Consistent Evidence Robust Recognition: Faithful Attribution Regularization under Geometric Transformations

链接: https://arxiv.org/abs/2607.23835
作者: Xianghao Jiao,Ruoyu Chen,Wei Wang,Jiazi Hu,Jiawei Liang,Shangquan Sun,Shiming Liu,Qunli Zhang,Xiaochun Cao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Attribution methods are widely used to characterize the evidence underlying model predictions, yet their potential to improve model behavior remains underexplored. Attribution inconsistency under label-preserving geometric transformations may indicate transformation-sensitive evidence reliance, motivating attribution regularization. However, such supervision is valid only when attribution faithfully reflects the evidence driving predictions. Existing self-supervised methods typically align gradient-based maps such as Grad-CAM, whose limited faithfulness means that attribution consistency need not imply consistency of the underlying decision process, leaving transformation robustness unresolved. We propose an annotation-free attribution regularization framework based on submodular search over image regions. By measuring how candidate subsets affect model outputs, the search extracts compact, class-discriminative evidence as search-derived supervision. We further introduce a submodular ranking loss with path-consistency and termination-alignment terms that respectively align spatially corresponding candidate rankings along paired search trajectories and encourage the transformed trajectory to satisfy the stopping criterion at the target terminal step. The loss provides a differentiable surrogate for regularizing both final attributions and the otherwise discrete evidence-selection process. Experiments on ImageNet-100 show that our method substantially improves attribution stability, Insertion, and Deletion on ViT-B/16 with only a 0.28-point accuracy drop, with similar gains on ViT-L/16. On ImageNet-1K, it improves transformed-input accuracy on ResNet-50 and ConvNeXt-B while limiting the clean-accuracy drop to 0.30 points, demonstrating more consistent evidence reliance with minimal performance loss. Code will be released soon.

[CV-88] Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms

链接: https://arxiv.org/abs/2607.23803
作者: Ruiqi Wu,Bingliang Jiao,Ruize Han,Hangzheng Yu,Xunkai Jiang,Shining Wang,Yuanqi Hu,Wenxuan Wang,Peng Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory reconstruction in multi-platform cooperative perception. Unlike conventional multiple object tracking, MvMoAT faces frequent viewpoint shifts that distort appearance and undermine cross-view association and temporal tracking. We propose FUSION, a viewpoint-robust Feature Unification framework for multi-view aSsociation and IdentificatiON. Its Multi-cue Adaptive Combination (MAC) module adaptively integrates viewpoint-invariant cues with appearance features to improve cross-view association, while Online Multi-view Feature Synchronization (OMFS) aggregates pedestrian features across historical and cross-view frames for temporally consistent tracking. We also introduce RealMvMoAT, a large-scale benchmark featuring substantial inter- and intra-camera viewpoint variation. It contains 504.9K frames from 7 cameras (5 UAV and 2 ground views) across 10 scenes, with over 7.3M identity-labeled bounding boxes. All cameras exhibit random and substantial motion. To the best of our knowledge, RealMvMoAT is the largest MvMoAT dataset to date. Its scale, viewpoint diversity, complex platform motion, and realistic trajectories provide a comprehensive resource for future research. Experiments on RealMvMoAT and six public benchmarks show that FUSION achieves state-of-the-art performance.

[CV-89] PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

链接: https://arxiv.org/abs/2607.23794
作者: Chi Phan,Tianyi Zhang,Yufeng Wu,Qiaochu Xue,Jiajie Zhang,Linghan Cai,Zeyu Liu,Sudong Wang,Yueming Jin,Dan Hu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification. However, existing pathology benchmarks and vision-language models (VLMs) are still largely developed under single-scale settings, limiting their ability to learn clinically meaningful multi-magnification reasoning. Moreover, naively constructed visual question answering (VQA) tasks may be susceptible to text-only or superficial visual shortcuts, leading to unreliable assessments of visual understanding. To address these limitations, we introduce a benchmark and training framework for shortcut-resistant cross-scale pathology reasoning. We design an Adversarial Text-only Screening strategy for semantic reasoning questions and a Structure-controlled Distractor Sampling strategy for visual grounding questions, encouraging models to rely on cross-scale visual evidence. Based on this pipeline, we construct PathScale-VQA, a high-quality cross-scale pathology VQA benchmark with 10,373 multiple-choice questions grounded in 1,368 diagnostic paths across multiple magnification levels. Building on the semantic reasoning set, PathScale-R1 is optimized through Difficulty-driven Reasoning Distillation supervised fine-tuning followed by reinforcement learning with a Scale-aware Reasoning Structure reward, which encourages the use of evidence across magnifications. Extensive experiments demonstrate state-of-the-art performance of PathScale-R1 on cross-scale reasoning tasks and effective transfer to conventional single-scale pathology VQA. Our code is available at this https URL.

[CV-90] RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction

链接: https://arxiv.org/abs/2607.23758
作者: Han Jiao,Chen Liu,Jiakai Sun,Zhanjie Zhang,Mengyuan Yang,Yimeng Li,Mofan Zhou,Kun Zhan,Lei Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-specialized optimization methods can produce high-quality road representations, but they typically require per-scene training and scene-dependent coverage design around the driving trajectory, limiting scalable reconstruction over newly collected roads. To address these limitations, we introduce RoadVGGT, a road-structure-aware feed-forward framework that reconstructs compact Gaussian road surfaces without test-time per-scene optimization. RoadVGGT uses a geometric foundation model to exploit multi-view images together with provided pose and depth observations, and predicts dense pixel-aligned Gaussian attributes through a learned Gaussian head. To make these dense predictions usable for large road surfaces, we align them into a consistent metric world coordinate system and fuse redundant Gaussians on the road-aligned XY plane through confidence-weighted grid fusion. Category-aware grouping and road–sidewalk junction protection further control fusion around vulnerable road structures. The resulting representation supports RGB and semantic bird’s-eye-view maps, elevation estimation, and novel view synthesis. RoadVGGT eliminates the need for per-scene optimization in prior methods, reconstructs complete road surfaces with a compact Gaussian representation, and improves image quality, semantic mapping, and elevation accuracy. Extensive experiments demonstrate the potential of geometric foundation models for scalable feed-forward road surface reconstruction.

[CV-91] DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation

链接: https://arxiv.org/abs/2607.23755
作者: Jianhan Lin,Yuchu Qin,Jiateng Yuan,Wenbo Zhang,Shuai Gao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ( t_rel ) of 1.31% and rotation error ( r_rel ) of 0.46 ^\circ . Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.

[CV-92] Source-Free Controlled Adaptation of Teachers for Continual Test-Time Adaptation

链接: https://arxiv.org/abs/2607.23735
作者: Anurag Roy,Riddhiman Moulick,Vinay Kumar Verma,Saptarshi Ghosh,Abir Das
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:In many real-world scenarios, encountering continual shifts in domain during inference is very common. Consequently, continual test-time adaptation (CTTA) techniques leveraging a teacher-student framework have gained prominence, allowing models to adapt continuously even after deployment. In such a framework, a weight-averaged mean teacher is used to produce pseudo-labels from test data for self-training. The mean teacher gets updated as an exponential moving average of the student parameters using a high value of momentum that is kept fixed even if different distributions of test data are encountered. To combat the resulting drift of the model, we propose a novel controlled teacher adaptation methodology that dynamically sets a proper momentum value depending on the quality of the incoming data. Additionally, we estimate class prototypes from the source pretrained model to help align the target data as they come in. Importantly, our method does not require access to source data or its statistics at any stage of the pipeline, making it truly source-free. We perform extensive experiments on benchmark datasets to demonstrate that our approach outperforms different state-of-the-art adaptation frameworks, many of which require access to source data.

[CV-93] LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratories

链接: https://arxiv.org/abs/2607.23704
作者: Haobo Wang,Baoli Sun,Anqi Zou,Dongsheng Huang,Zelin Lv,Ning Wang,Rui Li,Dongzhan Zhou,Weiyu Guo,Zhihui Wang,Wanli Ouyang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Under review. Haobo Wang and Baoli Sun contributed equally. Code and data: this https URL

点击查看摘要

Abstract:The deployment of embodied agents in self-driving laboratories could accelerate scientific discovery, yet their reliability is constrained by the irreversible and safety-critical nature of chemical experiments. Progress is further hindered by scarce failure data and the lack of fine-grained evaluation protocols. To address these challenges, we introduce LabRobFail, a failure-centric framework for learning and evaluating robotic failure analysis in chemical laboratories. LabRobFail-Sim injects controllable failures at the control, physics, and semantic levels, enabling the construction of LabRobFail-Data, which contains over 20,000 trajectories across 70+ task scenarios, five failure categories, and 11 fine-grained failure types. LabRobFail-Bench evaluates six capabilities spanning task understanding, failure detection, temporal localization, severity assessment, failure classification, and actionable correction. We further develop LabRobFail-VLM, a domain-specialized vision-language model that generates structured failure diagnoses and recovery instructions. On seen environments, it achieves 92.58% failure-detection accuracy and 85.58% temporal-localization accuracy, substantially outperforming general-purpose VLMs. When integrated as a real-time supervisor, it improves downstream VLA task success rates by 10-20 percentage points, demonstrating the value of fine-grained failure understanding for closed-loop recovery and reliable laboratory autonomy. Our code and data are available at this https URL

[CV-94] Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation MICCAI

链接: https://arxiv.org/abs/2607.23694
作者: Changjing Liu,Yiming Huang,Beilei Cui,Liangjing Shao,Long Bai,Yanheng Li,Haoxuan Che,Hongliang Ren
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by The 2nd MICCAI Workshop on Efficient Medical AI

点击查看摘要

Abstract:Efficient surgical segmentation empowers clinical diagnosis, intraoperative monitoring, and downstream robotic pipelines for reconstruction and simulation. Although prompt-driven foundation models like Segment Anything Model 3 (SAM3) achieve strong segmentation performance on natural images, surgical data exhibits domain gaps against its pre-training data, resulting in degraded segmentation accuracy. Furthermore, existing medical SAM methods require full-parameter fine-tuning, incurring heavy computational consumption and low efficiency. To address these limitations, this work proposes a parameter-efficient Low-Rank Adaptation (LoRA) adaptation of SAM3 for surgical concept segmentation. We inject low-rank adapters into the prompt encoder, detector and tracker while fully freezing the vision backbone, which only optimizes 0.98% of the total model parameters and supports training on a single consumer GPU. Comprehensive experiments demonstrate that our method consistently outperforms zero-shot SAM3 and other mainstream baselines, and the generated segmentation results can be directly deployed to support downstream robotic surgical scene reconstruction and physical simulation pipelines.

[CV-95] GNM Head: A Generative aNthropometric Model of the human head

链接: https://arxiv.org/abs/2607.23687
作者: Stylianos Ploumpis,Jan Bednarik,Gaspard Zoss,Ruslan Guseinov,Luca Prasso,Prashanth Chandran,Oliver Boyne,Vasileios Choutas,Timo Bolkart,Daoye Wang,Menglei Chai,Di Qiu,Sebastian Winberg,Gilles Rainer,Lewis Bridgeman,Delio Vicini,Jérémy Riviere,Yannick Boetzel,Alexander Koumis,Jay Busch,Cynthia Herrera,Jacob Still,Scott Ysebert,Peter Lincoln,Sergio Orts Escolano,Christoph Rhemann,Erroll Wood,Thabo Beeler,Stefanos Zafeiriou
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: The GNM is publicly available at: this https URL

点击查看摘要

Abstract:Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large vision models, allowing for tight spatial control of generated imagery. However, existing publicly available models are typically limited in anatomical scope, modeling only outer geometry while ignoring intra-oral and ocular structures, and frequently suffer from reduced geometric quality stemming from low-fidelity input datasets. In this report we introduce a new parametric model dubbed Generative aNthropometric Model (GNM), named as a homophone of the human genome. GNM encompasses the head, face, neck, eyeballs, teeth, and tongue, and it is built on an extensive database of high-resolution 3D scans combined with high-quality anatomy specific artist-made samples. This report details the data provenance, the model architecture including the specialized sub-models for the ocular and intra-oral structures, and shows its SotA performance on fitting target 3D face scans. To foster community innovation, the complete GNM framework is made publicly available.

[CV-96] Perturbation-Aware Diffusion-Guided Hybrid Segmentation for Robust and Annotation-Efficient Plant Stress Phenotyping

链接: https://arxiv.org/abs/2607.23680
作者: Gurbhit Chaurakoti,Soumyashree Kar
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Semantic segmentation in agricultural imagery is often evaluated under in-domain protocols, yet practical deployment requires robustness to appearance perturbations, limited annotations, and cross domain shift. This paper presents a diffusion-guided hybrid segmentation framework in which U-Net, DeepLabV3+, and SegFormer backbones generate coarse masks that are refined by Denoising Diffusion Probabilistic Models (DDPM), latent diffusion, or semantic-guided diffusion. The framework is evaluated through a 3x3 architectural screening study on PlantSegV3, followed by boundary-constrained optimization, perturbation-guided retraining, low-data evaluation, constrained hyperparameter screening, and controlled cross-domain adaptation. On PlantSegV3, the best selected hybrid model achieves 71.83% refined mean Intersection-over-Union (mIoU) and 26.10% refined Boundary-F1, and the selected models remain stable under substantially reduced supervision, demonstrating strong annotation efficiency. Perturbation analysis identifies grayscale conversion, fog, coarse dropout, and shadow as the most disruptive appearance shifts, and the resulting augmentation policy substantially improves robustness during retraining. The adapted models further show effective transfer to external agricultural datasets under limited target supervision, indicating that diffusion refinement and boundary-aware optimization provide transferable structural priors. Overall, the results show that carefully matched backbone-refiner pairings, combined with perturbation-aware retraining, can improve structural delineation and robustness under realistic resource and distribution constraints.

[CV-97] Contrastive Parameter Disentanglement for Multi-modal Remote Sensing Image Generation

链接: https://arxiv.org/abs/2607.23673
作者: Yu Zhang,Wenda Zhao,Haojun Tang,Haipeng Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Existing remote sensing image generation methods are largely confined to single-modality synthesis and therefore fail to exploit the complementary information inherent in multimodal imagery. To address this limitation, we propose a contrastive parameter disentanglement framework for multimodal remote sensing image generation, which generates semantically consistent and structurally aligned images across multiple modalities, including optical, infrared, and synthetic aperture radar (SAR), from a single text prompt. Specifically, we introduce a contrastive parameter disentanglement module that disentangles shared semantics from modality-specific attributes at the parameter level within an orthogonal core subspace. Based on this module, we develop a disentangled optimization strategy that first constrains the parameter matrix A of the LoRA adapter to capture modality-invariant semantics through a multimodal contrastive objective and then guides multiple parameter matrices B to learn modality-specific attributes under text conditioning. This strategy enables the simultaneous generation of multimodal images with consistent semantic content and distinct modality characteristics. Furthermore, to ensure structural alignment across the generated images, we devise a query-key structure transfer mechanism that jointly models multimodal sampling trajectories during inference by transferring structural correlation priors from an anchor modality to the remaining modalities. Extensive experiments demonstrate that our method outperforms state-of-the-art remote sensing image generation approaches in terms of generation quality, semantic consistency, and structural alignment, while also achieving superior performance in the downstream object classification task.

[CV-98] RRTrack: Robust and Recoverable Object 6D Pose Tracking for Dynamic Scenes

链接: https://arxiv.org/abs/2607.23669
作者: Junyue Li,Ye Zheng,Yifan Chen,Zhe Sun,Xuelong Li
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 6D Pose Tracking, 10 pages, real-world experiments, training-free

点击查看摘要

Abstract:Robust object 6D pose tracking is critical for robotic systems operating in dynamic and occluded scenes. Per-frame estimators are accurate but computationally expensive, while current trackers struggle with fast motion and complete occlusion due to their reliance on continuous visibility. To address these challenges, we present RRTrack, an efficient, recoverable object 6D pose tracker that enables robust tracking through fast motion and target disappearance–reappearance. RRTrack introduces a 2D–6D closed-loop tracking strategy that integrates memory-based video object segmentation (VOS) with 6D pose refinement. The 2D branch maintains target localization, and the 6D branch verifies geometric consistency before memory updates. In addition, a DINOv2-based dual-bank template matching module is developed to recover lost targets by jointly exploiting offline synthetic templates and online observation anchors while maintaining real-time efficiency. We also introduce a synthetic RGB-D benchmark comprising three robotic scenarios with fast motion and full occlusion. Experimental results on the synthetic benchmark demonstrate that RRTrack improves equal-subset mean ADD-S AR by 66.3% and ADD-S AUC by 65.7% over FoundationPose while achieving 55.2 FPS. Real-world experiments further validate the robustness of RRTrack under noisy sensing conditions. Project page: this https URL

[CV-99] XMatchAD: A Cross-Modal Matching Perspective on Reconstruction-based Anomaly Detection

链接: https://arxiv.org/abs/2607.23658
作者: Mingxiu Cai,Zhe Zhang,Gaochang Wu,Tianyou Chai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: accepted by IEEE Transactions on Image Processing

点击查看摘要

Abstract:The remarkable success of reconstruction-based methods in Unsupervised Anomaly Detection (UAD) lies in their ability to identify and localize anomalies by modeling discrepancies between input images and their reconstructed counterparts. However, these approaches often struggle to capture subtle anomalies and tend to produce blurred anomaly boundaries, which significantly limits their effectiveness, particularly in complex multi-class scenarios. To address these issues, we present XMatchAD, a novel UAD framework that reinterprets the task from a pseudo cross-modal matching perspective. Specifically, the input and reconstructed images are treated as two complementary modalities and their matching relationships are precisely exploited for anomaly detection. First, a pre-trained feature extractor is employed to encode discriminative representations. Second, an attention-guided cross-modal matching mechanism is introduced to match local inter-modal anomaly-related patterns while mutually refining the features. This enhances the sensitivity to anomalies with diverse shapes and subtle deviations and significantly improves the precision of anomaly detection and localization. Third, we design an adaptive frequency-aware fusion module that further delineates sharp anomaly boundaries through the coupling of high-frequency components from cross-modal multi-scale representations. Comprehensive evaluations on MVTec-AD, VisA, and MPDD benchmarks demonstrate that our method consistently achieves superior performance, outperforming state-of-the-art methods in multi-class anomaly detection and localization. The code will be released at this https URL.

[CV-100] GRAPE: Graduated Routing for Articulated Portrait mesh Estimation

链接: https://arxiv.org/abs/2607.23657
作者: Yunfei Liu,Lijian Lin,Ye Zhu,Yu Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Articulated portrait mesh estimation is fundamental to 3D understanding, avatar generation, and immersive interaction. Existing approaches primarily rely on 3D Morphable Models (3DMMs). However, face-centric models suffer from the “floating head” assumption, conflating head pose with global rotation due to the lack of neck kinematics. Conversely, body-centric models lack high-fidelity facial expression capabilities. Furthermore, current methods struggle to disentangle jaw articulation from expression blendshapes, often over-relying on expressions for mouth opening. These limitations make monocular portrait recovery difficult across representation, supervision, and anatomical parameter estimation. To address these limitations, we introduce GRAPE(Graduated Routing for Articulated Portrait mesh Estimation). We build a Portrait Parametric Model (PPM) with an explicit torso-to-head kinematic chain and a canonical injection step to merge FLAME and the SMPL-X torso. We propose a Progressive Anatomical Alignment (PAA) network, which is composed of a pretrained portrait encoder, a Graduated-Mask Router, and coarse-to-fine experts that follow the portrait anatomical prior. We then train this network with multi-source supervision that combines sparse anatomical keypoints, feature distillation, foreground mask constraints, and relative geometry constraints. Experiments show that GRAPE improves portrait mesh recovery quality, pose alignment, and jaw–expression disentanglement over prior methods. We also demonstrate that our method can benefit the downstream tasks of audio-driven talking-head generation and 3D portrait generation.

[CV-101] WGDnet: Wishart-guided Geometric-aware Deep Network for PolSAR Image Classification

链接: https://arxiv.org/abs/2607.23638
作者: Junfei Shi,Haojia Zhang,Yu Cheng,Yuke Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Polarimetric Synthetic Aperture Radar (PolSAR) classification underpins all-weather Earth observation. Conventional Wishart methods depend on rigid handcrafted operators with limited adaptability, while mainstream deep networks ignore PolSAR native Wishart scattering statistics. Additionally, fixed convolution windows fail to capture multi-scale, multi-directional terrain patterns, harming boundary detection and small-object characterization. To mitigate these drawbacks, we propose WGDNet, a Wishart-guided geometric-aware deep network. It integrates three core designs: (1) learnable Wishart convolutions with directional kernels for multi-scale statistical edge feature extraction; (2) an orientation-prior aggregation module that estimates dominant local directions and confidences to refine directional Wishart outputs adaptively; (3) GAnet, a scale-direction adaptive geometric-aware convolution that dynamically reshapes sampling grids to model anisotropic terrain and retain fine details. Our contributions lie in learnable Wishart statistical modeling, orientation-prior feature aggregation, and geometry-adaptive convolution. Evaluations across four real PolSAR datasets verify WGDNet surpasses existing state-of-the-art approaches in classification accuracy and boundary fidelity.

[CV-102] PathSelect: Sequential Token Selection for Whole Slide Pathology

链接: https://arxiv.org/abs/2607.23631
作者: Jingzhi Chen,Landi He,Zehong Chen,Peihang Wu,Lijian Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominantly rely on spatial sampling or training-free pruning, which risk diluting weak but informative signals, leading to the loss of critical diagnostic evidence due to the spatially diffuse nature of pathological cues. We reformulate WSI token pruning as a sequential selection process, enabling the model to autonomously learn an optimal routing strategy rather than relying on static heuristics. We herein propose a decoupled routing framework integrated as an active plugin into the fully pre-trained SlideChat base model, leaving both the slide encoder and large language model frozen. To provide continuous gradients for the non-differentiable pruning operation during training, we introduce PathSelect. PathSelect employs a variance-preserving noise gate to modulate each patch’s information flow via a differentiable Soft Top-K operator, paired with a diagonal-attention Denoiser that recovers the perturbed representations without semantic leakage. At inference, the PathSelect module is entirely detached. Relying solely on the trained Scorer, a deterministic Hard Top-K operator executes adaptive, data-dependent trajectory termination, significantly accelerating downstream generative processing with exceptionally low sequential token selection latency. Driven by an empirical average of only 44.86 tokens under a maximum constraint of K = 128, our framework achieves 74.00% overall accuracy on SlideBench (TCGA), representing an approximate 36.6x spatial token reduction relative to the uncompressed baseline average while consistently outperforming sampling-based counterparts.

[CV-103] Markerless Motion Capture in Routine Clinical Upper Limb Assessments: Validity and Insights Beyond Ordinal Scoring

链接: https://arxiv.org/abs/2607.23608
作者: Tim Unger,Olivier Lambercy,Roger Gassert,Andreas R. Luft,R. James Cotton,Chris Easthope Awai
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:The Action Research Arm Test (ARAT) is a widely-used upper limb outcome measure in neurorehabilitation, but its ordinal scoring is subjective and suffers from limited sensitivity and specificity. We evaluated whether artificial-intelligence (AI)-based markerless motion capture (MMC), embedded into ARAT assessments during clinical routine, accurately reconstructs upper limb movement and yields valid, objective kinematic metrics carrying clinically meaningful information beyond the ordinal score. Across 47 sessions from 20 mixed-neurological patients (1,174 ARAT tasks), biomechanical reconstruction was accurate and robust across impairment levels, and kinematic metrics showed the discrimination pattern expected of a construct-valid measure. In longitudinal case studies, the metrics added the specificity and sensitivity the ordinal score lacks: a domain decomposition exposed patient-specific recovery profiles underlying equal ARAT gains (specificity), and kinematic improvement continued to be detected after the ARAT had saturated (sensitivity). MMC in clinical routine can thus provide valid, objective, sensitive, and specific kinematic measurement complementing ordinal scoring.

[CV-104] ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion

链接: https://arxiv.org/abs/2607.23600
作者: Guo Yurong,He Yufei,Li Yonghao,Chang Dongliang,Zhang Ke,Ma Zhanyu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 5 figures

点击查看摘要

Abstract:Controllable infrared-visible image fusion aims to integrate complementary thermal and structural information with flexible region-aware modulation, producing fused images that adapt to diverse user requirements and downstream tasks. However, existing methods typically rely on predefined discrete control conditions, leading to a sparse space that fails to support fine-grained modulation demands. To address this, we propose ConFusion, a novel framework that learns the continuous fusion space via Gaussian-conditioned spatial-aware modulation, enabling instance-level fine-grained controllable infrared and visible image fusion. ConFusion employs a dual-branch architecture to disentangle modality-invariant and modality-specific representations under joint reconstruction and text-guided semantic alignment. Gaussian-conditioned instance modulation variables coupled with Grounded SAM-based instance masks guide instance-level fine-grained modulation through the Mask-Guided Specific Feature Modulator, while the Text-Driven Invariant Feature Enhancer improves semantic consistency and enhances fusion. During inference, the multimodal large language model parses user intents into instance-level modulation variables to guide image fusion. Extensive experiments show that ConFusion achieves state-of-the-art performance across multiple metrics in both fusion quality and downstream tasks, while supporting fine-grained controllable image fusion. Our code is available at this https URL

[CV-105] Weakly Supervised Instance-Level Gleason Pattern Estimation Using Primary and Secondary Labels MICCAI

链接: https://arxiv.org/abs/2607.23594
作者: Nao Sugeta,Kaito Shiku,Shinnosuke Matsuo,Ryoma Bise
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to MICCAI workshop 2026 (AMAI)

点击查看摘要

Abstract:In prostate cancer histopathology, the Gleason Score is determined by the most frequent (Primary) and second most frequent (Secondary) Gleason patterns within a whole-slide image. Although these slide-level labels are routinely available in clinical practice, instance-level Gleason annotations are rarely provided, making patch-level learning challenging. We propose a Multiple Instance Learning (MIL) framework that estimates instance-level Gleason patterns from slide-level Primary and Secondary labels. The proposed method formulates instance-level learning according to the clinical definition of the Gleason Score by aggregating instance predictions into class counts and explicitly modeling the Primary pattern, Secondary pattern, and their dominance. Experimental results demonstrate that the proposed formulation enables effective instance-level learning and outperforms existing MIL approaches on the SICAP-MIL dataset.

[CV-106] JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents WWW

链接: https://arxiv.org/abs/2607.23588
作者: Yunlong Lin,Zixu Lin,Zhaohu Xing,Biqiang Li,Chenxin Li,Haonan Wang,Haitao Wu,Hengyu Liu,Jianghai Chen,Kaituo Feng,Kaixin Li,Shawn Chen,Shijue Huang,Sixiang Chen,Tsung-Yi Ho,Wenxuan Huang,Xiangyan Liu,Xiaomeng Hu,Xuanhua He,Yan Sun,Yunqing Zhao,Zhiqin Yang,Zehan Wang,Zhengyang Tang,Tianyu Pang,Xiangyu Yue
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 9 figures. Project page: this https URL Code github: this https URL

点击查看摘要

Abstract:Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent’s external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.

[CV-107] SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion

链接: https://arxiv.org/abs/2607.23580
作者: Kavish Jhaveri,Arya Shah
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages, 4 tables, 4 figures

点击查看摘要

Abstract:Existing vector-sketch models treat recognition and generation as separate tasks, leaving a gap for streaming interfaces that must understand a drawing as it is being made. We present SketchMamba, a single causal sequence model that continuously classifies a sketch from any partial prefix while simultaneously generating its continuation. We achieve this by applying a dense per-step classification loss to a selective state-space backbone. Evaluated on a 58-class subset of the Quick, Draw! dataset, SketchMamba yields 94.93% final-step accuracy and a progressive-accuracy Area Under the Curve (AUC) of 0.706, crossing 90% of its final accuracy by the time 70% of the strokes are drawn. In a matched-budget comparison, the 1.55 million-parameter backbone ties a causal Transformer while outperforming recurrent and convolutional baselines. Ablations confirm that the dense supervision regime, rather than the architecture alone, drives the early-prediction capability. The results demonstrate that a single causal hidden state can unify progressive recognition and autoregressive generation without auxiliary encoders or task-specific branching.

[CV-108] Neuromorphic Object Detection: An In-Depth Study and Future Directions

链接: https://arxiv.org/abs/2607.23576
作者: Jianing Li,Dianze Li,Arren Glover,Xiaopeng Fan,Guoqi Li,Chiara Bartolozzi,Ryad B. Benosman,Yonghong Tian
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by Proceedings of the IEEE

点击查看摘要

Abstract:Conventional frame-based cameras face significant challenges in detecting objects under high-speed motion blur or in low-light environments. Neuromorphic cameras provide asynchronous visual streams with high temporal resolution and a wide dynamic range, offering a promising solution for object detection under challenging conditions. Despite the development of numerous models and the emergence of various applications in neuromorphic object detection, there is still a lack of deep understanding and standardized benchmarks to assess progress and address key challenges. In this paper, we provide a comprehensive survey and benchmark of existing neuromorphic object detection algorithms. Specifically, we first present a problem description, review the available datasets, and revisit the evaluation metrics. We then explore existing neuromorphic object detection approaches from various perspectives, including event representation, temporal modeling, multimodal fusion, asynchronous processing, low-latency processing, and energy-efficient computing. Furthermore, we evaluate a wide range of representative neuromorphic object detection models and offer detailed analyses of the comparative results. Finally, we discuss unresolved issues in neuromorphic object detection and propose potential future research directions. We hope this survey and benchmark will be a valuable resource for researchers and provide guidance for future advancements in neuromorphic object detection.

[CV-109] D3O: Dynamic Distribution Distillation for Ordinal Regression

链接: https://arxiv.org/abs/2607.23575
作者: Chunlai Dong,Yaojun Hu,Yuyang Xu,Haochao Ying,Jian Wu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 10 pages, 5 figures, ACMMM2026

点击查看摘要

Abstract:Ordinal regression is widely used in scenarios where labels are discrete yet inherently ordered. In practice, however, ordinal labels are often obtained by discretizing underlying continuous semantics through subjective human judgment, resulting in ambiguous boundaries and annotation noise. Such uncertainty challenges existing methods that rely on fixed supervision targets, which may reinforce biased ordering under subjective annotations. To address this limitation, we propose D3O, a dynamic distribution distillation framework that replaces static supervision with training-driven evolution of ordinal label distributions via self-distillation. Specifically, we introduce a contrastive ordinal-aware label enhancement module that leverages vision-language alignment to recover refined label distributions capturing both inter-class ambiguity and instance-level uncertainty. Furthermore, we design a CDF-based cross-layer interaction distillation mechanism to propagate cumulative ordinal structure across network hierarchy, ensuring consistent ordinal geometry in intermediate representations. Extensive experiments on four general ordinal regression tasks demonstrate that our proposed D3O consistently outperforms existing approaches, particularly under severe class imbalance and noisy supervision. These results highlight the effectiveness of dynamic supervision in learning robust ordinal representations beyond fixed targets. The code will be publicly available.

[CV-110] GaitFace: A Multimodal Dataset for Long-Range Person Identification

链接: https://arxiv.org/abs/2607.23542
作者: Alain Komaty,Luis S. Luevano,Vidit Vidit,Anjith George,Zeina Al Amine,Sébastien Marcel
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted in IJCB 2026 , see this https URL

点击查看摘要

Abstract:Efficient border control is becoming a significant global challenge, mainly due to severe congestion and extended passenger waiting times. To mitigate these bottlenecks and facilitate passenger flow, biometric technologies are increasingly deployed to streamline identity verification and enhance crossing efficiency. Technical limitations frequently impede biometric identification, particularly in long-range surveillance, where systems must deal with adverse atmospheric conditions and degraded image quality. While high-quality frameworks like BRIAR exist, they are frequently restricted to specific government agencies. This paper introduces GaitFace, a new public dataset that contains face and gait data captured at long distances. To ensure that the research reflects authentic border scenarios, we use Pre-Enrollment data, where a traveler registers via a mobile device, and “In-the-Wild” captures, which records individuals at a distance across multiple viewing angles and different cameras. Benchmarking SOTA face and gait models reveals that current architectures fail under low-resolution and elevated viewpoints despite success with optical assistance. GaitFace exposes these critical vulnerabilities, providing a rigorous public benchmark to drive more robust, unconstrained biometric research.

[CV-111] Geometry Meets Semantics: Fractional Gradient Stabilization for Semantic-Driven Bounding Box Optimization in Visual Detection Tasks

链接: https://arxiv.org/abs/2607.23530
作者: Qi Ming,Haitian Yang,Xudong Zhao,Mingjing Zhao,Liuqian Wang,Nanqing Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Bounding boxes are fundamental for object localization in visual detection tasks. Among them, oriented bounding boxes are widely used in visual detection tasks, which provide a more precise directional representation. Generally, IoU-based losses are widely adopted to optimize box regression. However, we observed that IoU-driven box optimization suffers from two key issues: (1) it relies solely on geometric properties while ignoring semantic cues; (2) orientation optimization suffers from unstable gradients, causing oscillations in orientation convergence. In this paper, we propose a Fractional Semantic IoU loss to achieve unified semantic-geometric learning with gradient stabilization. First, we design a semantic similarity metric to guide IoU optimization, building a Semantic IoU loss (SIoU loss) with an adaptive gradient gating mechanism. Then, we revisit the gradient instability issue in oriented box optimization and extend the SIoU loss to a fractional-order formulation to build the \textbfFractional \textbfSemantic \textbfIoU \textbfloss (FrSIoU loss). The FrSIoU loss accumulates historical IoU states to regularize abnormal gradients during bounding box optimization process. Extensive experiments demonstrate that our approach achieves stable performance gains across different bounding box formulations and diverse visual detection tasks. The code will be available on GitHub.

[CV-112] ATCNet-CIAM for Multi-Session Motor Imagery EEG Signal Classification

链接: https://arxiv.org/abs/2607.23522
作者: Le Huu Son Hai,Nguyen Chi Hai,Truong Viet Vu,Nguyen Phuc Nguyen,Nguyen Thai Anh,Ngo Hoang Tu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: This manuscript has been accepted for publication in the International Conference on Intelligence Systems and Robotics for Sustainable Development (ISRSD) 2026

点击查看摘要

Abstract:Motor imagery (MI)-based electroencephalography is widely used in non-invasive brain–computer interfaces (BCIs), but robust decoding remains challenging due to inter-subject variability and cross-session non-stationarity. This work proposes ATCNet-CIAM, an enhanced attention temporal convolutional network that integrates a lightweight channel-integrated attention module (CIAM) into the ATCNet framework to improve channel-spatial feature representation for MI decoding. The proposed model is evaluated on BCI Competition IV-2a, BCI Competition IV-2b, and the multi-day WBCIC-MI dataset under standard, within-session, and cross-session protocols. Experimental results show that ATCNet-CIAM achieves 86.32% accuracy on BCI IV-2a and 87.96% on BCI IV-2b under the standard protocol, while reaching 89.46% and 83.64% in the within-session WBCIC-MI on 2C and 3C, respectively. The proposed framework consistently improves classification stability and robustness under session-varying conditions, and ablation study confirms the complementary contribution of the proposed architectural components.

[CV-113] Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction

链接: https://arxiv.org/abs/2607.23517
作者: Chaonan Ji,Jinwei Qi,Peng Zhang,Bang Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:We present a real-time human-centric world model for upper-body interactive generation, aiming to synthesize coherent local world dynamics centered on a person, where coordinated body, hand, and facial motions evolve jointly with controllable human-object discrete interaction. To this end, we adopt a continuous-discrete joint control scheme with two complementary components: a continuous human state and a discrete interaction state. For continuous human-state control, we introduce a unified implicit representation based on multi-scale motion encoding, in which motion latents from the upper body, hands, and face are fused into a shared latent space. This multi-scale design improves expressiveness across different spatial scales, captures fine-grained human dynamics more effectively, and enables direct control without explicit retargeting. For discrete object interaction-state control, we represent object contact using a small set of language-encoded discrete interaction states, where text serves as an explicit interaction-state command, such as \emphno contact or \emphgrasp, rather than an open-ended generation prompt, and we further construct a dedicated rendering pipeline for human-object interaction data to supervise such discrete interaction states. By combining continuous implicit human-state control with discrete interaction-state control, our model enables precise modeling of how a person moves and interacts with the local environment, including controllable changes to nearby scene states. Finally, we distill the model for efficient streaming real-time inference, achieving 25 FPS on two H100 GPUs. Experiments demonstrate improved fine-grained motion fidelity, more realistic hand-object coordination, and effective real-time interaction, establishing a practical step beyond motion reproduction toward real-time human-centric world modeling.

[CV-114] MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

链接: https://arxiv.org/abs/2607.23511
作者: Zhijing Cheng,Xuancheng Zhang,Donglin Di,Lei Fan,Baorui Ma,Hao Li,Xun Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downstream planner predicts trajectories conditioned on this context. We argue that this one-way perception-to-planning interface forces sensor inputs into a compact representation, losing the fine-grained details critical for planning. Moreover, by constraining the planner to this compressed context, it is difficult to leverage the rich representations offered by modern vision foundation models. To address these issues, we propose MOJITO, a unified sensor-to-action framework for end-to-end autonomous driving built on modal joint learning. MOJITO removes the cascaded interface and instead performs block-wise Modal Joint Attention that simultaneously updates action, image, and LiDAR features, allowing the planner to directly access multi-modal features during action generation. MOJITO achieves 88.9 PDMS on the NAVSIM v1 dataset and 88.4 EPDMS on the more challenging NAVSIM v2 dataset, setting a new state-of-the-art. Extensive experiments further demonstrate strong scalability, instruction following, and diverse trajectory generation. Code and models are available at this https URL.

[CV-115] MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation

链接: https://arxiv.org/abs/2607.23504
作者: Yuqi Liu,Shengju Qian,Tianyuan Qu,Mingxian Lin,Zixuan Wang,Xin Wang,Bei Yu,Jiaya Jia
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to maintain long-horizon visual history for trajectory consistency while executing actions with low latency. Existing video-based VLN approaches typically struggle to satisfy both demands simultaneously. To address these challenges, we propose MemVLN, a novel VLN framework that achieves state-of-the-art performance with real-time inference efficiency (14 FPS). MemVLN utilizes a visual encoder to process continuous observations and a Large Language Model (LLM) to interpret instructions and generate actions. Central to our approach is an Episodic Memory management that applies pyramidal resolutions. This mechanism concentrates computation on immediate percepts while retaining compressed long-term history. Complementing to this design, we introduce Procedural Memory for fast action with a compact vocabulary of atomic mid-level actions to bypass auto-regressive decoding latency. Experiments on VLN-CE show that MemVLN-4B surpasses the baseline Qwen3-VL-4B architecture by 5.8% SR in R2R and 9.7% SR in RxR, while achieving a 7 \times speedup in inference latency.

[CV-116] oken-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

链接: https://arxiv.org/abs/2607.23493
作者: Musa Tur Farazi,Nufayer Jahan Reza
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age. However, classifying internet memes remains computationally challenging due to the intricate interplay between visual cues and embedded, often stylized, text, particularly in low-resource languages like Bengali Language. This paper addresses the detection of political intent in Bengali memes by introducing Multimodal Cross-Attention Fusion framework. We first leverage a Vision-Language Model to extract high-fidelity OCR text from noisy meme images. Subsequently, we encode visual and textual features and synthesize them through a cross-modal multi-head attention mechanism that aligns semantic tokens with visual regions. We also investigate the integration of a domain-specific political lexicon as a knowledge prior. Experimental evaluation on the PoliMemeDecode1 dataset shows that our attention-based fusion significantly outperforms unimodal baselines and standard concatenation methods, achieving a state-of-the-art Macro-F1 of approximately 0.94. Interpretability analyzes further confirm that the model effectively learns to ground textual semantics in visual evidence.

[CV-117] o Erase or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion

链接: https://arxiv.org/abs/2607.23492
作者: Shaswati Saha,Rajasekhar Anguluri,Manas Gaur
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Concept erasure techniques (CETs) edit text-to-image diffusion models to erase undesired targets such as NSFW content or copyrighted styles, while preserving model utility on benign concepts. Current CETs face a trade-off between erasure robustness and utility: stronger edits erase the target more reliably but degrade utility on non-target concepts, and vice versa. This stems from how existing methods define what to erase and what to preserve. Many CETs rely on static concept banks specified manually, generated by LLMs, or selected by CLIP image-text similarity. Such banks do not model how prompts steer the model during denoising, leaving it vulnerable to triggers that reintroduce the target while suppressing nearby benign concepts. We present Preservation-aware Adaptive Ranked Subspace Expansion (PARSE), a training-free framework for robust concept erasure in latent diffusion models. Given a target, PARSE queries the diffusion model with classifier-free guidance to dynamically discover target-inducing erase concepts and nearby retain concepts in the model vocabulary. It then edits the cross-attention value space with a preservation-aware projection that removes target directions while leaving retain directions intact. For triggers beyond this vocabulary-indexed space, PARSE iteratively searches for re-emergence triggers by textual inversion and adaptively expands the erased subspace only when a new trigger direction does not conflict with retain semantics. We also introduce the Balanced Erasure Utility Score (BEUS), which combines robustness (ASR under multiple attacks) and utility preservation (FID) via bounded monotone transforms and harmonic mean aggregation. Experiments on NSFW, artistic style, and object erasure, with a large-scale robustness-utility analysis over many CET baselines, show that PARSE erases multiple concepts robustly without sacrificing post-edit utility.

[CV-118] Learning Sampling Parameters for Diffusion Models

链接: https://arxiv.org/abs/2607.23488
作者: Arisrei Lim,Yossi Gandelsman
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: Code will be released

点击查看摘要

Abstract:Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then held fixed across prompts and denoising timesteps, even though different prompts and stages of generation can benefit from different parameter values. We introduce LeSAMP, a framework for learning prompt-conditioned, timestep-varying sampling parameters. We formulate parameter selection as a reinforcement learning problem: Given a user prompt, a large language model is trained to emit schedules for the chosen sampling parameters. We optimize our model using rewards from human preference models and VLM-as-a-judge. We evaluate our model on Flux.1 [dev] and Stable Diffusion 3.5, and find that compared to baselines, LeSAMP has a win rate of up to 68.12% using human preference scores and 73.37% using VLM-as-a-judge. These gains are validated in a user study where we achieve win rates of up to 59.46% over previous baselines. Our results suggest that learned sampling-parameter policies provide a complementary approach to existing post-training methods for improving diffusion model outputs.

[CV-119] VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

链接: https://arxiv.org/abs/2607.23472
作者: Tianxiao Chen,Hanmo Chen,Huajin Chen,Bo Li,Qi Ye,Peng-Tao Jiang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: tech report

点击查看摘要

Abstract:Modern video generation models can synthesize visually compelling and temporally coherent clips, yet controlling their physical behavior remains difficult with standard text and image conditions. The core challenge is a conditioning bottleneck: material response, contact interaction, deformation, and motion trajectory are continuous and relational physical cues that are hard to specify exhaustively in language but can be demonstrated naturally by video. We propose VIPER, a Visual In-Context Physics Reasoning framework for reference-guided image-to-video generation. Given a target image, a brief target prompt, and a reference video, VIPER treats the reference as a dense visual demonstration of the desired physical process rather than an appearance template. It uses a Multimodal Large Language Model (MLLM) to extract reference-derived physical cues and guide a pretrained image-to-video generator through a hierarchical training strategy, enabling physical behavior transfer while preserving the visual prior of the base generator. To support this setting, we construct VIPER-19K, a curated dataset with material, trajectory, and physical-impact annotations, together with filtered reference-target pairs. Experiments on an unseen validation set show that VIPER achieves stronger reference-video physical similarity and higher human preference than representative video generation and video-as-prompt baselines, while maintaining competitive general video quality. Qualitative results further demonstrate that VIPER can transfer reference-derived physical behavior to new target scenes without requiring carefully engineered prompts.

[CV-120] Robust 6-DoF Object Pose Tracking with Built-In Recovery under Occlusions and Rapid Object Motions

链接: https://arxiv.org/abs/2607.23468
作者: Balázs Opra,Léo Ghafari,Thomas Stewart,Cyrill Stachniss
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 4 figures. Accepted for publication in IEEE Robotics and Automation Letters (RA-L)

点击查看摘要

Abstract:Real-time 6-DoF object pose tracking is essential for many robotics applications, and several approaches exist. Yet even today’s approaches remain unreliable under temporary full occlusions and rapid object motions. Once tracking is lost, most methods struggle to detect the failure and recover automatically, often requiring manual re-initialization. In this paper, we address the problem of robust model-based 6-DoF tracking of unseen objects from RGB-D data, especially in scenarios with occlusion and fast motion. We propose a novel method that combines efficient learning-based keypoint matching with optimization-based alignment and introduces a novel failure detection and recovery module. Our system monitors pose reliability, detects tracking divergence or occlusions, and performs a global re-detection and pose estimation step that robustly verifies recovery candidates before resuming tracking. Our evaluation on standard tracking benchmarks and on a new dataset of occluded and fast-moving scenes shows that our method matches state-of-the-art accuracy on easy tracking sequences, maintains high tracking speed at 57.6 frames per second, and provides the most robust tracking performance under challenging conditions. Thus, we believe that our approach is a relevant step forward in robust 6-DoF object tracking from RGB-D data.

[CV-121] Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance

链接: https://arxiv.org/abs/2607.23451
作者: Weixiang Zhou,Jiabei Zuo,Yuhao Wang,Cong Wang,Huchuan Lu,Zhixun Su
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by IEEE TIP 2026. The version of record may differ slightly

点击查看摘要

Abstract:Multi-modal object Re-Identification (ReID) aims to retrieve specific objects by integrating complementary information from multiple modalities. However, existing multi-modal ReID methods do not effectively address background interference suppression or achieve tri-modal alignment, instead focusing on pairwise feature fusion. Moreover, many current aggregation approaches suffer from high computational complexity. To address these limitations, we propose PRISM, a novel multi-modal ReID framework built upon Prompt-S6 (PS6) and semantic-aware knowledge guidance. PS6 maintains the linear complexity and strong sequence modeling capability of Mamba while enabling efficient cross-modal interaction. Leveraging these advantages, we design two key components: Semantic-Driven Token Pruning (SDTP) and Progressive Fusion Network (PFN). Parsing semantic priors from the segmentation foundation models, the SDTP then leverages these priors and applies dynamic token pruning to suppress background noise and refine feature representations. The PFN progressively aggregates multi-modal features to achieve tri-modal alignment and fully exploit modality complementarity. With the proposed modules, PRISM generates more robust multi-modal representations under complex scenarios. Extensive experiments on four multi-modal object ReID benchmarks demonstrate the effectiveness and efficiency of our approach. The source code is available at this https URL.

[CV-122] Semantic Semi-Incremental Data-Association-Free Object SLAM

链接: https://arxiv.org/abs/2607.23384
作者: Yihao Zhang,Jungseok Hong,John J. Leonard
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Data association between landmark measurements and landmark variables has long been a central challenge in SLAM, as estimation accuracy depends critically on associating measurements with the correct landmark variables. Recent advances in deep learning have created new opportunities for the problem; data association can now leverage not only positional measurements but also semantic information about object landmarks, such as class labels from neural object detectors and feature vectors from visual foundation models. In this paper, we present a generalized data-association-free SLAM framework that jointly estimates data associations, robot poses, landmark positions, and landmark semantics from odometry, and positional and semantic measurements of landmarks. The proposed framework (i) creates a synergy between data association and landmark semantics estimation; (ii) adopts a semi-incremental estimation scheme for improved accuracy and computational efficiency; and (iii) provides a principled justification, guidelines, and heuristics for landmark-number estimation, improving the interpretability and practical usability of the framework. The proposed framework and algorithms are evaluated on synthetic and real-world datasets with two types of semantic information, class labels and real-valued feature vectors, and demonstrate superior performance compared to strong baselines.

[CV-123] UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models ECCV2026

链接: https://arxiv.org/abs/2607.23373
作者: Ioannis Maniadis Metaxas,Adrian Bulat,Alberto Baldrati,Anestis Zaganidis,Yassine Ouali,Hyeonuk Kim,Georgios Tzimiropoulos
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ECCV 2026

点击查看摘要

Abstract:Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction or smaller language models, the vision encoder is largely overlooked, typically deployed as a monolithic, computationally heavy feature extractor. Moreover, there is no previous effort that designs a vision encoder for LVLMs directly optimized for on-device latency. In this paper, we present UltraViT, a vision encoder for LVLMs, explicitly designed and optimized for on-device performance. Specifically, by taking into account real on-device latencies, we systematically design a pyramidal architecture that strategically integrates and adapts heterogeneous spatial mixers at the macro-block level. Furthermore, to pre-train UltraViT, we propose a novel two-stage generative pre-training strategy: cultivating rich spatial features via dense distillation, followed by direct generative supervision from a capacity-mixed frozen LLM. Compared to standard contrastive and SSL, we show that our pre-training is much more effective for achieving high-level semantic grounding for UltraViT needed for the subsequent generative multimodal alignment of LVLM training. Extensive experiments demonstrate that our on-device latency-informed design combined with our tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.

[CV-124] he Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models

链接: https://arxiv.org/abs/2607.23335
作者: Moshiur Farazi,Sameera Ramasinghe,Bekir Sait Ciftler,Mahbub Ahmed Turza,Shafin Rahman
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that the optimiser almost always decides on zero: across five injection designs, every gated pathway becomes behaviourally closed, with accuracy invariant to ablating the pathway at inference even when the gate parameter would nominally pass 30-45% of the signal. We attribute this suppression phenomenon to two regimes, a dead-gradient regime formalised through the caption-invariance of image-derived signals, and a negative-utility regime in which the auxiliary signal actively hurts the loss. Rather than fight suppression, we exploit it: we regularise LoRA fine-tuning with geometric auxiliary losses from hyperbolic visual relational graphs (IoA-driven entailment cones and angular repulsion on the Lorentz manifold), coupled only through the forward pass at training time and dropped at inference. Disaggregating GQA by question type exposes a clean dissociation. Three configurations without geometric losses at inference lose 2.85-3.39pp on relational questions while gaining ~1pp on attribute questions; a fourth that trains with the losses but infers through a soft prompt loses 5.14pp on rel for only +0.23pp on attr, so training-time regularisation alone does not protect relational accuracy without a geometric inference pathway. Configurations that keep the geometric pathway at inference preserve vanilla-level relational accuracy and match the attribute gain. Out of distribution on VSR, the RMS-prefix recipe preserves the spatial signal; stripping the geometric losses (G2) collapses VSR by 4.6pp, isolating them as the OOD source. A secondary result: embedding-norm alignment is necessary for generation-safe prefix injection, and learnable gates should be replaced with fixed, non-optional injection at matched scales.

[CV-125] Hiding in Plain Sight: An Effective Physical Adversarial Patch Attack against Visual-Infrared Fused Face Detection

链接: https://arxiv.org/abs/2607.23292
作者: Qiucheng Yu,Tao Ni,Yihe Zhou,Jiayimei Wang,Qingchuan Zhao
类目: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deep learning-based visual-infrared fused face detection models are increasingly deployed across a wide range of applications, yet they remain susceptible to adversarial patch attacks. Most prior attacks target either the visual or the infrared image alone in the digital domain, which renders them ineffective against fused models in the physical world. Moreover, many of these methods are readily noticeable, as their patch patterns deviate substantially from those seen in the real world. In this paper, we introduce VIPatch (Visual-Infrared Patch), a novel physical adversarial patch attack that produces inconspicuous, realistic, and natural-looking patches for facial images. Specifically, VIPatch crafts a gradient-color mask together with a band-aid sticker across both the visual and infrared images, and jointly optimizes these two elements; the resulting digital patches further guide the fabrication of their physical counterparts. Experimental results show that VIPatch achieves competitive attack success rates (over 90%) in both the digital and physical domains, while keeping the patches unobtrusive to human observers.

[CV-126] What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features

链接: https://arxiv.org/abs/2607.23271
作者: Chen-Yi Lu,Yueh-Shao Chen,Somali Chaterji
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., “a dog” vs. “not a dog”) to nearly identical embeddings, rendering them insensitive to negation. We attribute this failure to a phenomenon we call Representational Collapse: by tracking compositional divergence and visual alignment across the CLIP text encoder, we show that middle layers build compositional syntax, but the final layers collapse this structure as visual alignment rises, producing a syntax-blind final representation. To recover the lost negation signal without altering pretrained weights, we propose PeakPatch, a lightweight post-hoc correction system that intercepts the encoder at its compositional peak. An Embedding Correction Network (ECN) uses cross-attention to extract a negation-specific signal from the peak layer, anchored to a stable baseline, and predicts a deviation vector that re-injects the lost syntax into the final-layer embedding space. A complementary Score Correction Network (SCN) predicts bounded scalar score offsets for discriminative tasks. Both modules are trained jointly end-to-end while all CLIP parameters remain frozen, adding only 5.2M parameters (3.5% of the backbone) and preserving the standard cosine similarity interface. On NegBench, PeakPatch achieves 74.3% on COCO MCQ (+35.1 over CLIP, +17.8 over the best encoder fine-tuning method) and 65.5% on VOC MCQ, while outperforming all fine-tuning baselines on fully out-of-distribution negation retrieval despite training only 3.5% of the parameters. The corrected embeddings also transfer to text-to-image generation (+18.4 negation score) and generalize across ViT-B/32, ViT-L/14, and SigLIP backbones. Project URL: this https URL.

[CV-127] WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

链接: https://arxiv.org/abs/2607.23265
作者: Yuhui Zeng,Wang Chen,Jinfa Huang,Tianyu Xie,Yongdong Luo,Jiayi Ji,Xiawu Zheng,jiebo Luo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages, 10 figures

点击查看摘要

Abstract:Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. While recent efficient methods attempt to compress tokens via hard pruning or uniform merging, they operate strictly in the spatial feature domain, where robust structural context and discriminative semantic details are inherently entangled. In this work, we propose WaveZip, a joint signal-frequency-domain framework for efficient video inference. Driven by the insight that temporal redundancy resides in low-pass approximation scales while spatial saliency strongly correlates with high-frequency components, WaveZip leverages Discrete Wavelet Transforms (DWT) to disentangle these signals. Temporally, it employs 1D DWT to analyze query-frame relevance, and the resulting high-frequency coefficients are further gated by inter-frame differences, with both signals jointly driving the dynamic allocation of a precise frame-level token budget. Spatially, a 2D DWT decomposes features into low-frequency approximations and high-frequency detail components, where the high-frequency coefficients are modulated within query-salient regions to regulate spatial reconstruction. Importantly, WaveZip requires no task-specific training and can be seamlessly integrated into off-the-shelf LVLMs to boost inference efficiency. Extensive experiments on long video understanding benchmarks demonstrate that WaveZip retains 99.6% of the full performance under an extreme 10x compression ratio, consistently outperforming state-of-the-art methods.

[CV-128] SARATR-X-v2: Scale-Aware Structural Pre-Training for SAR Foundation Models

链接: https://arxiv.org/abs/2607.23238
作者: Weijie Li,Yafei Song,Yongxiang Liu,Bowen Peng,Jie Zhou,Jingyuan Xia,Wei Yang,Tianpeng Liu,Zhen Liu,Li Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Masked image modeling has become a dominant paradigm for SAR pre-training, yet the design of the reconstruction target remains fundamentally unsettled. This article argues that a SAR pre-training target should satisfy two conditions to produce transferable representations: (i) physics-grounded stability, i.e., approximate invariance of the target operator to multiplicative speckle inherent in coherent imaging; and (ii) semantic scale compatibility, i.e., coverage of the heterogeneous spatial scales that downstream tasks demand. These two conditions are individually achievable but jointly difficult: physics-grounded stability favors fixed operators, while semantic scale compatibility favors data-driven composition. To this end, SARATR-X-v2 reconciles both within a single design. The target is constructed through fixed structural extractors spanning six receptive fields, from blind-spot local aggregation to directional log-ratio region contrast, and fused via learnable weights into one unified supervision signal for masked reconstruction. On twelve SAR benchmarks across classification, detection, and segmentation, SARATR-X-v2 achieves state-of-the-art transfer performance. Under synthetic speckle variation, the proposed target reduces perturbation drift in the learned representation by nearly two orders of magnitude relative to pixel-space supervision. Taken together, these results establish physics-grounded stability and semantic scale compatibility as a principled framework for pre-training target design under coherent imaging, and suggest that effective SAR pre-training is not about reconstructing more signal, but about reconstructing the right structural target.

[CV-129] A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

链接: https://arxiv.org/abs/2607.23235
作者: Zhijiang Tang,Jiaxin Qi,Kaihua Tang,Yuhua Zheng,Jianqiang Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22 pages. Code: this https URL

点击查看摘要

Abstract:Image captioning is a primary task in vision–language research, yet assessing how faithfully a caption preserves image semantics without relying on reference captions remains unsettled. Prevailing evaluations rely on human-annotated references, whose content reflects annotator intent and captioning proficiency. In this paper, we study a reconstruction-based principle for caption evaluation: a caption is as good as its capacity to enable reconstruction of the original image. However, because captioning inherently compresses visual information, it is impossible to recover all details, and pixel-wise comparison between reconstructed and source images is neither feasible nor meaningful. Through our in-depth analysis of the nature of captions, whose fundamental purpose is to transmit the semantic content of an image, we propose a revised principle: a caption is as good as its capacity to enable a reconstruction that is semantically equivalent to the original. To assess semantic equivalence, we test whether the reconstruction matches the original image across a suite of downstream vision–language tasks, yielding a reference-free, task-conditioned caption score. We characterize component-dependent limitations and introduce the lower-cost Captioning Turing Test Dataset (CTTD) surrogate.

[CV-130] BoneAgeTW2: Automated Skeletal Maturation Assessment via the Tanner-Whitehouse 2 Method Deep Learning and Clinical Report Generation with Distribution Curves

链接: https://arxiv.org/abs/2607.23224
作者: Juan Manuel Castillo Pinto
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Thesis research article. Code available at this https URL

点击查看摘要

Abstract:We present BoneAgeTW2, the first fully open-source system to automate the complete Tanner-Whitehouse 2 (TW2) clinical protocol for skeletal maturity assessment end-to-end. The system employs YOLOv8 for precise detection and localization of the 20 TW2 hand bones from radiographic images, and an EfficientNet-B3 backbone with 20 independent classification heads to assign maturation stages (A-I) to each bone simultaneously. From these predictions, the system automatically generates clinical PDF reports including interactive Gaussian distribution curves for all 20 bones, enabling direct comparison with population norms. The model is trained on the public RSNA Pediatric Bone Age Challenge dataset (12,611 hand radiographs) using a pseudo-labeling strategy to derive per-bone stage labels from global bone age annotations. The full codebase is publicly available at this https URL.

[CV-131] Counterfactual Motion Reliability Learning for Robust UAV Tracking

链接: https://arxiv.org/abs/2607.23209
作者: Yuehai Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Infrared unmanned aerial vehicle (UAV) tracking is challenging because the target is often small, low-contrast, and easily confused with thermal distractors or cluttered backgrounds. Recent Transformer-based trackers have achieved promising performance by learning strong appearance representations, but their responses can still be dominated by background structures when the target appearance is weak or ambiguous. A natural solution is to introduce temporal motion cues. However, in infrared UAV tracking, motion cues are not always reliable: camera jitter, dynamic backgrounds, sensor noise, and target disappearance may produce temporal variations that are stronger than the true target motion. Therefore, the key challenge is not simply how to use motion, but how to distinguish target-consistent motion from background-induced pseudo motion. To this end, we propose CMRTrack, a counterfactual motion reliability learning framework for robust infrared UAV tracking. CMRTrack first extracts temporal evidence from adjacent search regions using a lightweight motion evidence encoder. During training, a counterfactual target-erased history branch is introduced to construct hard motion references, encouraging the motion encoder to learn reliable target-consistent motion rather than arbitrary temporal changes. The learned motion evidence is then incorporated into a one-stream tracking framework through motion-guided token modulation and reliability-aware score fusion, enabling adaptive feature enhancement and response refinement. Extensive experiments on Anti-UAV410 demonstrate that CMRTrack consistently outperforms representative state-of-the-art trackers and significantly improves the OSTrack baseline, with ablation studies and qualitative analysis verifying the effectiveness of the proposed counterfactual motion reliability learning.

[CV-132] Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix

链接: https://arxiv.org/abs/2607.23194
作者: Zobeir Raisi,John Zelek
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 39 pages, 10 tables, includes supplementary material (folded in as Appendix S1-S4). Code: this https URL

点击查看摘要

Abstract:Scene Text Recognition (STR) models are trained almost exclusively on word crops of at most 25 characters, yet real deployments (signage, product labels, dense captions) require reading much longer text. This paper diagnoses that failure and then closes it. The diagnosis separates out-of-length failure into two simultaneously extrapolating axes (the encoder’s width axis and the decoder’s time axis) and shows that encoder width, not decoder length, is the dominant failure mode. Representation-side fixes bring only partial relief: training-free rotary rescalings recover at most 2-4 points of character error rate (CER), and a weighted fine-tuning recipe recovers 6-8 points while improving standard-benchmark accuracy, yet word accuracy on the Long Text Benchmark (LTB) stays near zero, because the residual gap lies in the decoding mechanism rather than the representation. We then close that gap at inference time, on an unmodified word-level checkpoint: the long image is sliced into overlapping crops at the model’s training width, each decoded independently and in-distribution, and the reads stitched by geometry-anchored edit-distance alignment. This procedure reaches 42.79-43.05% bucket-average word accuracy on LTB across two base checkpoints, matching the published state of the art (41.57%) and beating it by 11-12 points on the hardest bucket, at wall-clock parity with plain decoding; applied unchanged to the public PARSeq checkpoint it reaches 47.11%. Once chunking is applied fine-tuning no longer helps: the decoding-side fix alone matches purpose-built architectures. We release the diagnosis harness and implementation.

[CV-133] OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

链接: https://arxiv.org/abs/2607.23193
作者: Jinsen Su,Yongdong Luo,Yuexiao Ma,Yibo Hu,Meiguang Jin,Xiaowu Zheng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different moments. This cross-modal salience mismatch makes unidirectional guidance prone to discarding answer-critical cues under aggressive compression. We propose OmniScope, a training-free token compression framework that uses the query as a shared semantic anchor while estimating relevance separately for audio and video. OmniScope allocates modality-specific token budgets, prunes visual tokens with an anchor-delta strategy that preserves both global context and temporal changes, and merges audio tokens within each second to reduce redundancy while maintaining temporal continuity. Across four audio-video benchmarks and two Qwen2.5-Omni model scales, OmniScope achieves the best average accuracy across all compression settings. At 25% overall token retention, it delivers up to 3.53x prefill speedup and more than 15% GPU memory reduction, with only a 0.35-point drop in average accuracy. These results suggest a simple design principle for OmniLLM inference: share the query across modalities, but not the salience estimates. The code is available at this https URL.

[CV-134] Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design

链接: https://arxiv.org/abs/2607.23189
作者: Shenghao Yang,Hongtao Zhang,Yuhan Yi,Zhihao Tang,Zihao Cui,Lian Wen,Han Yan,Yuan Gao,Mingbo Zhao
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry. However, 3D garment generation remains in its nascent stage, where in the realm of fashion, the semantic information of diverse design elements exhibits intricate coupling relationships in 3D representations, posing substantial challenges for generating diverse 3D garments. In this work, to handle the above problem, We introduce Fashion-3DLR, a novel 3D garment generation framework that utilizes diverse design elements to create high-quality, versatile 3D garment assets. Specifically, to bridge the semantic gaps between different fashion elements, we propose a Garment Feature Fusion Diffusion Transformer (GFF-DiT) module to integrate 2D fashion design elements, e.g., sketch and texture, into latent space. Within the latent space, we then employ a rectified flow transformer to generate geometry latents, which can be decoded into various 3D garment representations, including 3D Gaussians and meshes. Furthermore, we integrate Fashion-3DLR into downstream tasks, achieving the 3D Gaussian Splatting (3DGS)-driven cloth physical simulation and mesh-based virtual try-on. Experimental results indicate that Fashion-3DLR surpass the previous state-of-the-art methods, which verify that the proposed work can generate well-structured, non-watertight garments capable of physical simulation and virtual try-on, underscoring its potential as a versatile 3D garment design tool.

[CV-135] A Scale-adaptive Vision Model Links C. elegans Neuronal Morphology to Behavior for Neurotoxicity Assessment

链接: https://arxiv.org/abs/2607.23183
作者: Haochao Ying,Shenchong Lv,Yutao Sun,Zijian Tu,Xufeng Jin,Yuyang Xu,Yizhe Wang,Wei Yang,Xiaomin Yue,Jian Wu,Peilin Yu
类目: Computational Engineering, Finance, and Science (cs.CE); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Neurological disorders are a leading cause of global disability and are increasingly linked to environmental chemical exposures. Yet neurotoxicity assessment still relies on hand-scored morphological readouts that are subjective and poorly predictive of behavioral outcomes. Caenorhabditis elegans provides a genetically tractable, 3R-compliant alternative, but quantifying neuronal phenotypes from confocal microscopy at scale remains computationally challenging: existing vision foundation models, trained on natural or radiological images, cannot resolve the sparse signals and multi-scale lesions of neuronal imaging. Here, we introduce a dedicated self-supervised vision model for C. elegans dopaminergic neurons, together with CeNeuMorph, a multi-grained confocal benchmark of 27,117 annotated images. Specifically, moving beyond standard Masked Autoencoders, we propose a scale-adaptive masked image modeling strategy that jointly learns representations across resolutions and patch sizes under a fixed token budget. By decoupling structural semantic learning from rigid grid constraints, the model effectively resolves the full spectrum of neurodegenerative lesions - ranging from fine dendritic beading to gross soma shrinkage - within a tractable computational framework. Finally, our model surpasses both generalist and biomedical foundation models across classification, segmentation and detection tasks. Fusing visual features with morphological descriptors enables prediction of dopamine-dependent behavioral deficits ( R^2=0.498 ). Screening 180 agrochemicals, we identify the benzimidazole moiety as a previously unrecognized determinant of dopaminergic neurotoxicity. Together, the work demonstrates how scale-adaptive self-supervised learning can connect morphology to function for a scalable alternative to mammalian in vivo models for neurotoxicity assessment and drug discovery.

[CV-136] owards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

链接: https://arxiv.org/abs/2607.23181
作者: Yihao Wu,Chenyi Xu,Liqi Yan,Chenhuan Cai,Geyong Min,Bin Lin,Fangli Guan,Jianhui Zhang,Pan Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening generalization and increasing computational burden. We propose BrainNav, a navigation framework grounded in the Principle of Minimal Sufficiency. BrainNav consists of three components: a Logical Anchor Model that implements instruction-aware selective perception to suppress environmental noise, a Minimalist Constraint Alignment module that serves as a compact cross-modal bottleneck, efficiently synchronizing discrete linguistic intent with continuous latent dynamics while filtering out redundant information, and a Compression World Model that predicts action-conditioned states within a condensed, low-rank latent space. These modules align semantic intent with spatial perception, enhancing the agent’s robustness and efficiency in complex tasks. Experiments show that BrainNav improves over prior SOTA by 2.0 % / 1.0 in SR/SPL on R2R-CE val-unseen and 0.94 % / 0.78 on RxR-CE val-unseen. These results indicate that minimally sufficient world representations provide an effective foundation for robust VLN.

[CV-137] CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

链接: https://arxiv.org/abs/2607.23159
作者: Shreshth Saini,Neil Birkbeck,Yilin Wang,Balu Adsumilli,Alan C. Bovik
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discarded. Training-free caching makes each rollout 2-3x faster at near-lossless quality. Composition is safe only if lossy caching preserves verifier rankings. We present the first study of whether caching corrupts candidate ranking in video test-time search. On Wan2.1-T2V-1.3B with an adaptive caching wrapper (~2x per-candidate speedup), ImageReward scores seed-matched cached and full rollouts. Median per-prompt Spearman rank correlation is 0.905, with 72% top-1 agreement on the VBench suite. VBench-2.0 replicates this result on a harder suite. Recomputing the cached winner at full compute retains 90-94% of the full-search gain. Errors cluster among near-tied candidates, making corruption self-limiting. This finding leads to CachedSearch. It explores every candidate with aggressive caching, then re-generates only the winner at full compute. At N=8, it captures 94.7% of best-of-N’s gain at 63% of the cost. Capture rises with width. At matched budget, it searches twice as wide for 38% more gain. The result holds from 1.3B-14B across six models and four families: Wan, LTX, CogVideoX, and Hunyuan. Wan2.1-14B matches the 1.3B model’s fidelity. Mid-trajectory pruning multiplies the exploration saving to 3.11x at 88.6% capture. Ports to other model families require recalibrating a single parameter, showing that fidelity tracks architecture rather than parameter count. CachedSearch is training-free, verifier-agnostic, and orthogonal to the search algorithm, making it a plug-in multiplier for test-time scaling.

[CV-138] DispatchRAG : Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video

链接: https://arxiv.org/abs/2607.23132
作者: Muhammad Sulthan Adhipradhana,Ehsan Javanmardi,Naren Bao,Manabu Tsukada
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Assessing the severity of a traffic accident scenario is important to decide which emergency service to dispatch. Missing an ambulance dispatch on a pedestrian accident is a fatal issue that can lead to death. Recently, Vision-Language Models (VLMs) have been a promising tool for accident reasoning, yet many VLMs are not grounded in real-life accident response protocols, making them not usable in accident severity assessment off-the-shelf. We introduced DispatchRAG, an accident assessor and dispatcher framework grounded in real-life Japanese traffic-accident response protocols, designed to enhance VLMs to generate an appropriate emergency response during an emergency scenario. Utilizing a RAG-based retrieval mechanism to retrieve the most relevant accident protocol and an LLM-powered reasoner to suggest the most proper response. To support evaluation, we introduce Accident Dispatch Dataset, a comprehensive dataset of accident assessment and emergency response according to Japanese accident response protocols adapted from the MM-AU dataset. We validate our framework on the Accident Dispatch Dataset, showing strong performance across various accident scenarios compared to the baseline VLM, pointing toward integration in autonomous vehicles that can automatically report both their own and nearby accidents.

[CV-139] Occlusion-Point Reuse for Ray-Traced Ambient Occlusion and Shadow

链接: https://arxiv.org/abs/2607.23122
作者: Haojie Jin,Fujia Su,Zehui Lin,Chenxiao Hu,Jierui Ren,Yuqing Yuan,Yanchen Zhang,Zhongtao Wang,Yisong Chen,Kangying Cai,Guoping Wang,Sheng Li
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Ambient occlusion (AO) and soft shadows are critical visibility cues for spatial perception in real-time rendering. Hardware ray tracing provides a direct way to evaluate these effects, enabling ray-traced AO and area-light shadows that avoid many limitations of screen-space AO and shadow mapping. However, real-time budgets allow only a few rays per pixel, leaving raw ray-traced estimates noisy and expensive. We present an occlusion-point reuse framework that reuses traced samples in the domain of first-hit occlusion points instead of directly reusing final shading values or light samples. This provides a ray-reuse formulation for AO, rather than merely filtering or reusing completed AO values. The key idea is to transform AO and area-light shadow estimators into occluder-domain integrals, then combine neighboring occluder samples with a multiple-importance-sampling (MIS) formulation. For both AO and shadows, we derive unbiased estimators that validate convergence to the transformed integrals, as well as biased estimators designed for practical real-time execution. The biased variants assume local first-hit occluder consistency; for shadows, this occluder-based assumption better matches local visibility geometry than the visibility-consistency assumption commonly used when reusing light samples. Experiments show higher AO and shadow quality than non-reuse ray-traced baselines, and better shadow quality than light-sample reuse at comparable cost.

[CV-140] SHReg: Strictly Rotation-Equivariant Point Cloud Registration via Spherical Harmonics

链接: https://arxiv.org/abs/2607.23096
作者: Chongjian Wang,Junjie Gao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Point cloud registration critically depends on local features that are both distinctive and robust to arbitrary 3D rotations. Existing learning-based methods typically approximate rotation invariance via fragile local reference frames or extensive data augmentation, providing only empirical invariance and often degrading under unseen rotational transformations. In this paper, we propose SHReg, a strictly rotation-equivariant point cloud registration framework grounded in the representation theory of SO(3) . By representing local geometric features as irreducible representations of SO(3) , SHReg guarantees exact equivariance under arbitrary rotations without relying on local reference frames. Built upon a spherical-harmonics-based equivariant backbone, SHReg jointly learns rotation-invariant descriptors for robust correspondence matching and rotation-equivariant features that preserve fine-grained orientation information. The preserved equivariant structure enables each correspondence to directly hypothesize a rigid transformation, reducing reliance on large-scale hypothesis sampling in conventional RANSAC-based pipelines and leading to improved robustness under challenging rotational variations. Extensive experiments on 3DMatch, 3DLoMatch, and KITTI demonstrate that SHReg consistently outperforms state-of-the-art methods in registration accuracy, particularly under large rotational perturbations.

[CV-141] Inverse Bayesian Inference for Extracting Lesion Dynamics from Longitudinal Spectral CT

链接: https://arxiv.org/abs/2607.23078
作者: Lukas Förner,Melina Wördehoff,Julian Steffens,Maximilian Schmutz,Rainer Claus,Josua Decker,Thomas Kröncke,Kartikay Tehlan,Thomas Wendler
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Longitudinal medical imaging captures temporal evolution of lesions, yet extracting the underlying dynamical parameters governing this evolution remains challenging. We propose an inverse Bayesian framework for inferring lesion dynamics from longitudinal spectral CT. We decompose spectral feature ( x ) evolution into three components: \beginequation* \fracdx_idt = A_i x_i + B \cdot n + C \cdot \Delta x_\textsat \endequation* where A_i captures intrinsic dynamics (lesion-autonomous evolution), B captures local environment tumour burden (organ tumour burden through satellite count coupling), and C captures environment/satellite state change (i.e., whether surrounding lesions move similarly or not). We demonstrate the framework on photon-counting NSCLC CT data from metastases, recovering distinct dynamical regimes: lung lesions exhibit significant satellite count coupling ( B=-0.34 , p0.05 ) suggesting competitive dynamics, while liver lesions show synergistic satellite behaviour coupling ( C\approx+1.0 , p0.05 ). Synthetic validation confirms parameter recovery, and cross-coupling analysis validates that our method detects non-zero coupling when present. This work establishes inverse dynamical inference as a principled methodology for extracting interpretable parameters from longitudinal imaging, moving beyond static feature extraction toward mechanistic characterisation of lesion behaviour. The code and data are available at: this https URL Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.23078 [cs.CV] (or arXiv:2607.23078v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2607.23078 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Lukas Förner [view email] [v1] Sat, 25 Jul 2026 07:02:15 UTC (1,915 KB)

[CV-142] DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding

链接: https://arxiv.org/abs/2607.23070
作者: Yilin Wang,Haochen Shi,Guanyu Chen,Weiqing Min,Jinkai Zheng,Chenggang Yan,Shuqiang Jiang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 8 figures

点击查看摘要

Abstract:Food segmentation is essential for applications such as intelligent catering, dietary assessment, and recommendation. However, existing benchmarks fail to capture the complexity of real-world dining scenes. The challenges of dense inter-dish overlap, fine-grained class similarity, and extreme long-tail class distributions exceed the fidelity of current datasets. To fill this gap, we introduce \textbfDishSeg24k, a large-scale dish-level segmentation benchmark with 24,096 images, 112,281 instances, and 278 fine-grained categories in real-world dining environments. Based on DishSeg24k, we further propose \textbfFood Expert-Adaptive Segmentation Transformers (FEAST) to address these challenges. FEAST models query-based decoding as a Markov Decision Process (MDP), where each decoder layer update is treated as a sequential decision step that explores uncertainty along dish boundaries. We further redesign the decoder with a reinforcement learning (RL)-guided Mixture-of-Experts (MoE) module, in which a dual-critic decoupled optimization scheme separates task-oriented query refinement from structure-aware expert routing. This design promotes expert specialization and prevents expert collapse under long-tail category distributions. Finally, extensive experiments on DishSeg24k demonstrate the state-of-the-art performance of FEAST, which outperforms previous methods by +3.21% mIoU, +3.68% mDice, and +4.00% mAcc, respectively. We further validate the effectiveness of FEAST on FoodSeg103. The dataset and code will be publicly released.

[CV-143] Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLM s ECCV2026

链接: https://arxiv.org/abs/2607.23046
作者: Jouwon Song,Woohyeong Kim,Kyeongbo Kong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ECCV 2026

点击查看摘要

Abstract:Recent high-resolution Multimodal Large Language Models (MLLMs) generate thousands of visual tokens per input, leading to a visual token explosion that introduces severe latency bottlenecks. While token pruning mitigates this issue, state-of-the-art subset-optimization methods typically rely on iterative subset construction to jointly capture visual diversity and instruction relevance. As visual token counts scale, this sequential dependency introduces significant selection overhead, severely limiting the translation of theoretical FLOPs reductions into actual wall-clock speedups. To address this limitation, we propose Single-Forward Pruner (SFPruner), a structural reformulation of visual token pruning that embeds redundancy control directly into the scoring space, bypassing the need for iterative combinatorial optimization. Our non-iterative framework achieves redundancy-aware importance selection in a single forward pass through two complementary mechanisms. First, to attenuate redundancy at the covariance level, we introduce a semantics-guided ridge leverage scheme. By integrating instruction relevance and visual saliency, this mechanism suppresses dominant covariance directions and mitigates representation bias. Second, ranking-based directional masking resolves residual overlap through asymmetric similarity competition, where higher-scoring tokens explicitly suppress redundant lower-scoring alternatives via parallel tensor operations. Extensive evaluations demonstrate that our approach maintains stable selection costs, reducing the token selection process by up to 110 ms, from 112.4 ms to just 2.5 ms at 512 tokens in Qwen2.5-VL. This structural efficiency successfully translates theoretical token reductions into tangible inference speedups while preserving highly competitive performance against state-of-the-art techniques under aggressive compression.

[CV-144] When Less Is More: A Controlled Benchmark of Lightweight CNNs for Satellite Land-Cover Segmentation on DeepGlobe

链接: https://arxiv.org/abs/2607.23024
作者: Atiq Ur Rehman,Joseph Michael Donovan
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Applications (stat.AP)
备注: 18 Figures, 8 Tables 33 Pages

点击查看摘要

Abstract:High-resolution satellite imagery is the backbone of good land-cover classification, and without that, environmental monitoring, urban planning, and sustainable resource management all fall short. Deep learning architectures perform well in semantic segmentation, but the efficiency-accuracy trade-off across classical convolutional encoders is not well quantified under controlled, reproducible conditions. This study compares five architectures VGG16, MobileNetV2, InceptionV3, AlexNet, and CNN on the DeepGlobe Land Cover Classification dataset using three progressively optimized iterations to isolate regularisation, transfer learning, and architectural depth. To ensure performance differentials reflect architectural properties, all experiments used identical preprocessing, hyperparameter, and training protocols without data augmentation or class-imbalance correction. At 24.98 MB, MobileNetV2_v1 had the highest overall accuracy (0.7906) and mean Intersection over Union (0.4625), outperforming deeper alternatives like InceptionV3_v2 (125.17 MB, accuracy 0.7610) and VGG16_v2 (71.13 MB, accuracy 0.7653). Class-wise analysis showed strength in urban, agricultural, and water categories, but rangeland-barren confusion showed that architectural optimization alone cannot optimize spectrally similar minority classes. Strong spatial generalization and crisp boundary delineation were confirmed on held-out test imagery, validating operational applicability. These results show that lightweight, transfer-learned models can match or outperform deeper models in resource-constrained remote-sensing environments, enabling scalable land-cover mapping.

[CV-145] OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

链接: https://arxiv.org/abs/2607.23023
作者: Quanyue Song,Yishan He,Yanbo Ding,Zhixiang He,Yongxiang Li,Caigui Jiang,Zhi Zhi Guo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for interactive avatar systems. However, extending these models to real-time interactive streaming remains challenging, as the generation horizon is unknown in advance and cross-modal identity consistency gradually degrades during long-term generation. To address these challenges, we propose OmniMate, a unified framework for open-ended real-time interactive audio-visual avatar generation. OmniMate jointly synthesizes visual content, speech, and audio effects in real time, enabling natural and immersive multi-turn interactions. To achieve adaptive response progression, we introduce a Generation Progress Controller (GPC) that explicitly models the generation progress of each streaming chunk, allowing the model to complete responses according to the desired progress and achieve seamless transitions between execution and listening states. To preserve long-term cross-modal identity consistency, we propose a Multi-Reference Conditioning Module (MRCM), which leverages multiple reference images and a reference speech segment to provide persistent visual and speaker identity cues throughout long-duration streaming interactions. Extensive experiments on an interaction-oriented adaptation of VerseBench demonstrate that OmniMate achieves high-quality, low-latency streaming generation while maintaining strong long-term audio-visual consistency. The results further show that OmniMate supports realistic, coherent, and responsive interactive avatar experiences over extended multi-turn conversations.

[CV-146] Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning ICML2026

链接: https://arxiv.org/abs/2607.22994
作者: Tao Zhang,Qixuan Fan,Yiyuan Liang,Yanjie Wang,Song Yan,Tian Tian,Jiahuan Zhou,Luxin Yan,Sheng Zhong,Xu Zou
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted by ICML 2026

点击查看摘要

Abstract:Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is effective, it raises concerns regarding privacy and storage. Thus, generative replay has emerged as a viable alternative, synthesizing old data using frozen pretrained text-to-image (T2I) models without any extra training. However, we observe that directly mixing synthetic old-class data with real new-class data during incremental training leads to significant performance degradation. This issue stems from a “domain shortcut”, where models rely on domain-discriminative features instead of semantic class cues. To address this, we propose DREAM ( \underline\mathbfD omain- \underline\mathbfR egularized \underline\mathbfE xemplar-free \underline\mathbfA lignment \underline\mathbfM odel), which uses a training-free generator to synthesize old-class data and eliminates domain shortcut via subspace rectification and orthogonal projection, while reinforcing semantic alignment through real-anchored prototype regularization. Extensive experiments on 4 datasets demonstrate that DREAM outperforms existing exemplar-free CIL methods and achieves state-of-the-art performance. Our source code is available at this https URL.

[CV-147] mmSimPrior: Learning Simulation Priors for Data-Efficient Real-World Generalizable Radar-Based Human Motion Reconstruction

链接: https://arxiv.org/abs/2607.22973
作者: Cheng Guo,Qiming Cao,Shengkai Xu,Haoyu Xie,Kaixiang Su,Pu Wang,Hongfei Xue
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 11 figures, including supplementary material. Project page: this https URL

点击查看摘要

Abstract:Millimeter-wave (mmWave) radar offers privacy-preserving and lighting-robust sensing for human motion reconstruction, but learning models that generalize across real deployments require diverse paired radar-motion data that are costly to collect. Simulation provides scalable supervision, yet models trained on clean synthetic signals transfer poorly because of multipath, clutter, response statistics, and resolution degradation. We present mmSimPrior, a simulation-pretrained framework that factorizes transferable knowledge into signal, motion, and radar-to-motion mapping priors. A multi-modal signal encoder is pretrained with a physics-informed domain-randomization curriculum that emulates propagation- and acquisition-level variations, while a joint-temporal tokenizer learns a discrete prior over plausible human motion. A shared mapping prior supports classification over a learned motion codebook for constrained zero-shot reconstruction and continuous regression for flexible limited-data adaptation. We further construct a 4.2M-frame, 31K-sequence dataset suite and introduce a No-Overlap Setting that excludes repeated complete subject-environment-location-motion configurations across adaptation and test. Experiments on mmSimPrior-Real and RT-Pose demonstrate consistent gains: with only 24 paired real sequences, mmSimPrior-Reg reduces MPJPE by 24.7% to 39.0% over the strongest baseline across the three environments, while mmSimPrior-Cls reduces zero-shot MPJPE by 8.5% without finetuning.

[CV-148] HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale

链接: https://arxiv.org/abs/2607.22959
作者: Aniket Sakpal,Yang Jiang,Rouzbeh Davoudi,Shayan Hassantabar,Mani Najmabadi
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 10 pages, 4 figures, 3 tables

点击查看摘要

Abstract:AI-generated video is increasingly used across marketing, product storytelling, and creative workflows, yet automated; high-precision quality control remains a major constraint to scaling production. We present HALLELUAI, an end-to-end system that moderates and regenerates image-to-video outputs to meet expert-level creative standards and deliver ultra-realistic videos with consistent end-user quality of experience (QoE) at scale. The system integrates a video moderation module that evaluates frame-level aesthetics, temporal motion fidelity, and fine-grained hallucination risks relative to the source image, with an agentic regeneration module that iteratively fixes failures through prompt refinement, controlled camera adjustments, targeted model or image switching, and structured retry strategies. The moderation logic is aligned with domain-specific creative guidelines and produces granular, machine-actionable feedback that directly drives regeneration. In human-in-the-loop evaluations with creative experts, HALLELUAI shows strong alignment and reliably outputs ultra-realistic, production-grade videos suitable for product and marketing placements at scale. This framework advances trustworthy AI generated video content by enforcing visual realism, brand safety, and strict input-image fidelity while enabling image-to-video generation at scale.

[CV-149] Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions ECCV2026

链接: https://arxiv.org/abs/2607.22931
作者: Quyen Tran,Hai Nguyen,Quan Dao,Zhuowei Li,Nam Le,Trung Le,Dimitris Metaxas
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: ECCV 2026

点击查看摘要

Abstract:Analytic Continual Learning (ACL) offers a computationally efficient alternative to gradient-based approaches. Recent ACL methods are based on Recursive Least Squares (RLS) and have achieved the state-of-the-art results compared to other alternatives. However, they falter significantly in Class-Incremental Learning scenarios characterized by Long-Tailed distributions. While the ill-conditioning of the autocorrelation (Gram) matrix is a known limitation of RLS, we demonstrate that class imbalance exacerbates this issue into a distinct spectral pathology: “tail” classes suffer from severe spectral collapse, rendering their subspaces numerically indistinguishable from noise. Standard Ridge Regression ( L_2 ) fails to address this effectively as it applies isotropic regularization - a uniform penalty that is insufficient to stabilize the tail without over-shrinking the head. To address this, we propose Geometry-Spectral Rectification (GSR), a theoretically grounded framework that treats long-tailed learning as a spectral regularization problem. Unlike standard isotropic regularization (Ridge) which uniformly penalizes all eigenvalues, GSR acts as an anisotropic spectral filter, selectively inflating the collapsed eigenvalues of tail classes. We construct a structured, data-dependent spectral perturbation matrix \Delta that selectively inflates collapsed tail eigen-directions of the Gram matrix. Theoretical analysis proves that GSR guarantees an improved stable rank for the Gram matrix, ensuring numerical stability. Extensive experiments show that GSR establishes a new state-of-the-art for analytic CIL, offering a superior trade-off between computational efficiency and robust generalization in long-tailed settings.

[CV-150] Layering Virtual Try-On ECCV2026

链接: https://arxiv.org/abs/2607.22924
作者: Chun Feng,Bowei Chen,Mengyi Shan,Ira Kemelmacher-Shlizerman
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ECCV 2026

点击查看摘要

Abstract:In the real world, fashion is about layering: adding a jacket over a shirt, or a sequence of adding and removing layers, rather than just a single-layer swap. This fundamental real-world task remains a challenge in existing Virtual Try-On (VTON) methods, which excel at single-layer replacement but are not designed to layer or de-layer an existing outfit. This paper proposes Layering Virtual Try-On (LVTON), a layering benchmark and method that preserves an existing outfit while enabling sequential layering. We find that current VTON paradigms are fundamentally ill-equipped for LVTON, as their reliance on cloth-agnostic representations and single-item datasets discards essential layering context. Our key insight is that the LVTON challenge must be disentangled into two distinct competencies: (1) General VTON Priors (e.g., deformation, identity preservation) and (2) Specific Layering Knowledge (e.g., layering order and occlusion reasoning). First, our model obtains general VTON priors by being trained on data produced by an automatic data generation pipeline that synthesizes samples from fashion videos via segmentation and inpainting. Second, the model is fine-tuned on a small, dedicated LVTON dataset to learn the layering logic. Our method achieves state-of-the-art results on our LVTON benchmark and demonstrates superior generalizability on traditional VTON benchmarks, setting new state-of-the-art results when fine-tuned and exhibiting zero-shot capabilities.

[CV-151] Controlling Embedding Spaces with Text-Conditioned Transformations ECCV2026

链接: https://arxiv.org/abs/2607.22919
作者: Joseph Fioresi,Fabian Caba Heilbron,Pankaj Nathani,Mubarak Shah,Kushal Kafle
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted at ECCV 2026

点击查看摘要

Abstract:Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification. These embeddings compress high-level semantics into a single vector, which comes at the cost of primarily expressing a dominant semantics like main object while suppressing other important attributes such as camera angle or color tone. We propose a text-conditioned transformation of visual embeddings that makes such attributes explicitly accessible. Given a natural language description of an attribute category (e.g., “color” or “art style”), a network generates an affine transformation that emphasizes the specified attribute. Conditioning on text enables it to learn many attributes simultaneously, accessing them at inference time through an intuitive interface. The network is trained to align transformed embeddings with the frozen latent space, enabling retrieval using existing large-scale embeddings without any re-encoding. When applied to a full set, the same mechanism transforms the latent space for attribute disentanglement tasks such as multi-clustering. By operating directly in latent space, our method provides a unified and efficient framework for controlling embedding spaces, demonstrating state-of-the-art performance across both attribute-based retrieval and multi-attribute organization tasks with near-zero inference cost. Project page: this https URL

[CV-152] Small-Pollinator Detection in Cluttered Field Video

链接: https://arxiv.org/abs/2607.22913
作者: Onur Onal(Iowa State University),Chen Chen(Institute of AI, University of Central Florida)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages, 2 figures. Code, experiment notebooks, and implementation details: this https URL

点击查看摘要

Abstract:Detecting pollinators in field video is challenging: targets are small, visually similar, and observed against cluttered vegetation under blur and occlusion. We present a systematic empirical study of small-pollinator detection under a practical single-GPU compute budget. Using the BuzzSpot challenge dataset, we compare YOLO and RF-DETR models across input resolutions and evaluate sliced inference, class-gated fusion, size-routed ensembling, and post-hoc temporal processing. RF-DETR Large at 1344-pixel resolution achieved our best hidden-test result, reaching 0.405 mAP50:95 and outperforming the 1120-pixel model (0.379) and the best single-model YOLO26m baseline (0.366). The strongest gains came from adopting RF-DETR and increasing its input resolution, indicating that detector choice and input resolution were more effective levers than added inference-time complexity; the resolution gain was strongest for small objects and the rarer bumblebee and moth classes. Sliced-inference fusion, size-routed ensembling, and warm-started 1536-pixel continuation did not surpass this result, while post-hoc temporal processing did not improve the leaked diagnostic evaluation. Error analysis identified bee-hoverfly discrimination as the clearest remaining bottleneck: neighboring frames rarely supplied correctly classified hoverfly evidence for post-hoc correction. These findings motivate learned feature-level temporal aggregation before the final classification decision.

[CV-153] AdaKAN: A dual-branch adaptive Kolmogorov-Arnold network for medical image segmentation

链接: https://arxiv.org/abs/2607.22891
作者: Dalia Alzu’bi,Deep Bhattacharyya,Ali Ayub,A. Ben Hamza
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Medical image segmentation is a fundamental task in computer-aided diagnosis, yet it remains challenging due to the complexity of anatomical structures and the variability across imaging modalities. In this paper, we propose AdaKAN, an Adaptive Kolmogorov-Arnold Network (KAN) that synergistically integrates convolutional operations with a novel efficient KAN (EffiKAN) block, comprised of an efficient attention mechanism and an adaptive KAN (AdaptKAN) module. This module features a dual-branch design: one branch employs a KAN layer with Bernstein polynomial activations for globally smooth and stable function approximation, while the other branch performs channel-wise refinement through projection operations and adaptive scaling. AdaKAN adopts a U-shaped architecture that effectively captures both long-range dependencies and fine-grained local features, overcoming the limitations of conventional convolutional and Transformer-based segmentation models. Skip connections are employed to preserve spatial details during encoding and facilitate accurate reconstruction during decoding. Extensive experiments conducted on diverse medical imaging datasets demonstrate that AdaKAN achieves state-of-the-art performance in segmentation accuracy.

[CV-154] Meshless Domain Randomization via Explicit Parameter Perturbation of 3D Gaussian Splatting

链接: https://arxiv.org/abs/2607.22890
作者: Felipe Nunes Carbone de Carvalho,Joyce de Morais Souza,Alan de Aguiar,Charles Morphy D. Santos,João Paulo Gois
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 10 pages, 4 figures

点击查看摘要

Abstract:Domain Randomization (DR) is a standard technique for closing the Sim-to-Real gap, yet traditional DR pipelines rely on classical computer graphics rendering driven by polygon meshes. For complex organic subjects, such as insect specimens, extracting and rendering textured meshes is challenging. To address this issue, we propose a meshless DR framework that operates on the parameter space of 3D Gaussian Splatting (3DGS). Our method employs two independent perturbation pipelines to synthesize randomized training datasets. First, a Photometric DR pipeline alters the baked illumination and color balance by modulating the Spherical Harmonics (SH) coefficients. Second, a Procedural DR pipeline isolates the subject’s geometric shape by replacing its original textures with 3D spatial noise. Finally, these perturbed radiance fields are composited over stochastically varied backgrounds using a rasterization engine. Our parameter manipulation provides a meshless alternative for generating robust datasets for complex geometries.

[CV-155] Same Predictions Different Reason s: The Effect of Quantization on Model Explanations

链接: https://arxiv.org/abs/2607.22872
作者: Kazi Kamruzzaman Rabbi,Md. Zami Al Zunaed Farabe,M. Sohel Rahman
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages, 3 Figures

点击查看摘要

Abstract:Post-training quantization (PTQ) has become a practical solution for deploying deep learning models on resource-constrained edge devices by compressing high-precision floating-point weights into low-precision representations without requiring retraining. Past research has demonstrated that quantization largely preserves classification accuracy; however, whether it also preserves the model’s internal reasoning remains an open question. This study presents a systematic evaluation on how static PTQ affects the interpretability / explainability of five widely used CNN architectures: VGG19, ResNet18, EfficientNet-B0, DenseNet161, and MobileNetV2 at INT8 and INT4 precision. We employ a dual interpretability framework that combines Grad-CAM for spatial attention analysis with LIME for input-level feature attribution, and systematically compare full-precision and quantized models on two binary classification datasets. Interpretability is evaluated using three complementary metrics: the Pearson correlation coefficient, structural similarity index, and top-20% IoU to capture distributional and structural variations in model explanations, supplemented by deletion/insertion faithfulness analysis. The results show that classification accuracy is not a reliable indicator of interpretability stability under reduced precision. DenseNet161 maintains strong feature consistency across both precision levels, whereas EfficientNet-B0, despite achieving competitive spatial attention and classification accuracy at INT8 precision, exhibits a substantial degradation in input-level feature attribution. These findings have direct implications for the trustworthy deployment of quantized models in applications with high interpretability requirements, demonstrating that architecture selection is as important as the quantization strategy.

[CV-156] Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

链接: https://arxiv.org/abs/2607.22864
作者: Patrick Rim,Tom Long,Ekta Prashnani,Ruth Rosenholtz,Ben Boudaoud,Peter Xenopoulos,Alex Wong,Joohwan Kim,Jae-Hyun Jung
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing benchmarks evaluate these models as black boxes, limiting their ability to identify the underlying causes of lower performance: when a model fails a spatial reasoning task, it remains difficult to ascertain whether the hurdle is perceptual, such as recognizing object boundaries, or cognitive, such as reasoning about occlusion to infer hidden geometry. We introduce Spatial-IQ, a hierarchical diagnostic framework that decomposes object counting in stacked 3D structures into 9 perceptual and cognitive sub-tasks organized by the developmental stages of human spatial cognition, with mental rotation as an additional target probe. Using NVIDIA Isaac Sim, we procedurally generated a diverse dataset of roughly 80,000 stacked 3D structures with per-task ground truth. We evaluate models across three output formats (free-response text, multiple-choice images, and image editing) alongside a human baseline. The Spatial-IQ framework shows that top-performing models often succeed at the target task (object counting) without succeeding on the lower-level sub-tasks intended to support it, and that models differ in how much of these hierarchical chains they preserve, often revealing shortcut behavior that raw target-task accuracy alone would obscure. Finally, we demonstrate that training models with chain-of-thought (CoT) supervision over our hierarchical sub-tasks, combined with reinforcement learning with verifiable rewards, significantly improves both spatial consistency across sub-tasks and target-task accuracy, supporting the value of the proposed decomposition as both a diagnostic tool and a training signal.

[CV-157] Robustifying pathology foundation models via fine-tuning

链接: https://arxiv.org/abs/2607.22861
作者: Alexandre Filiot,Oskar Thaeter,Benoit Schmauch,Lionel Guillou
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 17 pages, 4 figures, 7 tables

点击查看摘要

Abstract:Pathology foundation models (FMs) produce powerful tile-level representations which remain sensitive to scanner and staining variability, undermining deployment across laboratories. We develop a novel fine-tuning recipe that improves the robustness of pathology FMs to acquisition factors. Applied to ten different FMs, our fine-tuning strategy consistently improves robustness for every model as well as downstream performance, with no observed trade-off. On average, it raises the PathoROB robustness index by 23% (from 0.72 to 0.87) and increases the overall cross-benchmark performance by 43% on Patho-Bench, HEST and THUNDER combined, with individual gains reaching up to 72% in robustness (Phikon-v2) and 76% in performance (Midnight-12k). We publicly release the fine-tuned versions of Phikon-v2 (Phaet) and Midnight-12k (Mascaret) at this https URL.

[CV-158] Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling ICPR

链接: https://arxiv.org/abs/2607.22847
作者: Yuqi Hou,Zhuo Chen,Han Hu,Je Woo Kim,Jianbo Jiao,Hyung Jin Chang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ICPR(The International Conference on Pattern Recognition), Eye Tracking Techniques, Applications and Challenges (ETTAC 2026)

点击查看摘要

Abstract:Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typically process individuals independently, treating gaze as an i.i.d. quantity or predicting social semantics in isolation. Recent multi-person methods attempt to address this but often treat social relations as rigid, post-hoc classifications decoupled from the gaze estimation process. This oversimplification fails to capture the nuanced nature of social intent, which acts as an underlying driver of gaze behavior rather than a secondary categorical output. We address these limitations by proposing ANCHOR, a target-centric paradigm designed to decode gaze-anchored social intent by modeling the joint distribution of visual attention and latent implicit relations. Our approach surfaces these dependencies as the latent structural scaffolding of gaze behavior. The architecture utilizes a relational attention mechanism to capture fine-grained interpersonal links, leveraging feature-wise modulation for efficient multi-person parsing from a single vision backbone. To stabilize the training of this coupled formulation, we implement an optimization synergy to resolve the inherent conflicts between spatial gaze accuracy and latent social reasoning. This approach ensures robust generalization by seeking stable, flat minima while simultaneously harmonizing competing task gradients. We validate our framework on an extended benchmark featuring dense multi-person annotations and novel social influence rankings. Our results demonstrate state-of-the-art performance and provide the first quantitative evidence that implicit social hierarchies can be robustly disentangled and learned directly from static gaze patterns.

[CV-159] ID-V2V: Identity-Preserving Video Restylization SIGGRAPH

链接: https://arxiv.org/abs/2607.22830
作者: Yuancheng Xu,Mingming He,Pablo Salamanca,Li Ma,Yash Kant,Emmett Steven,Paul Debevec,Ning Yu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to SIGGRAPH Asia 2026

点击查看摘要

Abstract:In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restylization, which propagates scene, lighting, and style changes specified by an edited keyframe across a source video, while preserving facial likeness and performance, including expressions, eye gaze, and lip synchronization. A key obstacle is the absence of paired training data, as identity-preserving restylized video pairs are rare in real-world settings. To address this, we propose a decoupling of source-grounded identity preservation and edit-driven video synthesis. Our key insight is that facial appearance and expression should remain invariant, with illumination being the primary permissible variation. We therefore cast identity preservation as a video relighting problem, while modeling visual edit propagation as controlled video synthesis guided by the edited keyframe. Building on this formulation, we introduce ID-V2V, a video-to-video generative framework integrating complementary control signals: relit facial regions and facial normal maps tightly constrain facial likeness and performance, while edited keyframes and depth sequences enable flexible and temporally coherent generation. This design enables constructing training pairs from a single video, eliminating the need for scarce paired data. Extensive experiments demonstrate that ID-V2V significantly outperforms existing methods in preserving facial likeness and fine-grained facial performance, supports both single- and multi-subject scenarios, and delivers high visual quality, highlighting its potential as a human-centric tool for real-world content production. The code is available at: this https URL.

[CV-160] Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution ICANN2026

链接: https://arxiv.org/abs/2607.22808
作者: Md. Ajwad Hossain
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR); Image and Video Processing (eess.IV)
备注: Peer-reviewed and accepted to the DLMMDD Challenge Workshop at the 35th International Conference on Artificial Neural Networks (ICANN 2026). 7 pages, 1 figures

点击查看摘要

Abstract:The rapid advancement of text-to-image (T2I) models has necessitated robust Synthetic Image Source Attribution (SIA) methodologies. A critical challenge in SIA is the distribution shift between pristine training images and real-world deployed images, which undergo unknown post-processing operations such as JPEG compression and blurring. In this work, proposed for the DLMMDD Challenge at ICANN 2026, we introduce a dual-branch ensemble framework fusing Semantic Deep Learning with Mathematical Forensic Feature Extraction. The semantic branch employs EfficientNet-B0 regularized with Exponential Moving Averaging (EMA) and Label Smoothing. The forensic branch extracts 126 mathematical features – including SVD spectral profiles and Local Binary Patterns – from high-pass noise residuals, compressed via Truncated SVD and classified with XGBoost. Evaluated on a dataset of 10 generators where 55% of the test set is degraded, our approach achieves a private leaderboard accuracy of 95.60%. Furthermore, the entire pipeline is highly computationally efficient, requiring no GPU acceleration and executing end-to-end on a standard CPU in under 6.5 hours, highlighting the practicality and scalability of mathematical forensics for real-world deployment.

[CV-161] StateAct: Program State before Pixels for Long-Horizon Computer-Use Agents

链接: https://arxiv.org/abs/2607.22798
作者: Yan Yang,Xiangru Jian,Ziyang Luo,Zirui Zhao,Yutong Dai,Ziji Shi,Hanshu Yan,Jun Hao Liew,Silvio Savarese,Junnan Li
类目: oftware Engineering (cs.SE); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task data. Different states can produce the same pixels, while code can inspect and modify that state directly. StateAct is a code-first, multi-agent harness built around this distinction. Its main agent works directly with program state by using code, while a dedicated GUI subagent handles screenshot-and-click interaction on the few subgoals that need it, just 28 of 108 tasks and 1.1% of main-agent steps. The same direct access to program state also supports verification: an independent finish gate double-checks the saved result for structural failures, e.g., output that is missing, unsaved, or written to the wrong path. To stay on track over hundreds of steps, the main agent hands subgoals to fresh subagents, keeping its own context focused. On OSWorld 2.0, StateAct lifts Claude Opus 4.8 from 20.6% to 26.9% on binary success, and from 54.8% to 61.6% on partial success, at ~ 9x lower cost per task than the same model driven by screenshots alone; a code-only variant with no GUI subagent reaches only 45.9% partial, below that screenshot-based baseline’s 54.8%. In general, grounding action, verification, and memory in state, what we call state-grounding, shifts the main bottleneck from perception toward reasoning: failures depend more on what the agent thinks than on what it sees.

[CV-162] Inter-Reflective Gaussian Splatting for Robust and Efficient Inverse Rendering

链接: https://arxiv.org/abs/2607.22780
作者: Chun Gu,Xiaofei Wei,Zixuan Zeng,Yuxuan Yao,Li Zhang
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Faithful inverse rendering requires visibility and indirect radiance to explain secondary illumination and inter-reflection, yet rasterization-oriented Gaussian representations do not naturally support the secondary-ray queries needed to recover them. We present IRGS++ (Inter-Reflective Gaussian Splatting), a unified robust and efficient Gaussian inverse rendering framework. During transport-aware optimization, IRGS++ employs differentiable 2D Gaussian ray tracing on surface-oriented Gaussian primitives to query visibility and indirect radiance on the fly and evaluate the full rendering equation for inter-reflective transport. This physical core makes Gaussian inverse rendering physically grounded beyond rasterized appearance modeling. To make this backbone useful beyond low-gloss dielectric scenes, the framework incorporates metallic-aware material modeling and robust reflective initialization for glossy, specular, and metallic materials. To make it practical, multiple importance sampling and denoising stabilize finite-sample rendering, while mesh-based secondary-attribute queries reduce the cost of relighting under novel illumination. Quantitative evaluations on low-gloss and glossy benchmarks show improved decomposition and relighting quality together with favorable quality–speed trade-offs under the reported configurations, while real-world studies illustrate plausible relighting under novel illumination.

[CV-163] Cheap Probes Predict Expensive Training in 3D-CT Vision–Language Models

链接: https://arxiv.org/abs/2607.22771
作者: Renjie Liang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Picking the frozen image encoder for a 3D~CT vision–language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates. There are several encoders, several ways to compress their tokens, and several token budgets, and the combinations grow fast. Comparing them the usual way means fine-tuning a large language model (LLM) on each combination, and running the whole sweep this way needs far more compute than most groups can spend. We ask whether a cheap probe on the encoder’s cached embeddings can stand in for that comparison. We build an image-grounded probing benchmark over (encoder \times compression) cells, with clinical attribute families and two validation gates, scale-sanity and probe-separability, that keep each attribute well-scaled and decodable. These gates are the main methodological contribution. On this benchmark we compare a range of read-out heads, and in a preliminary study we pair each probe with its matched downstream task. The early signal is encouraging: the cheap probe orders the candidates in close agreement with expensive fine-tuning, at about r\approx0.95 on the cells measured so far. We read this as an ordinal claim, a ranking predictor rather than an exact estimate, and we are explicit about where it stays preliminary. If it holds up, encoder and compression choices can be screened in minutes with frozen-token probes, with full training spent only on the finalists.

[CV-164] Dementia Etiology Diagnosis via Collaborative Meta Knowledge Enhancement

链接: https://arxiv.org/abs/2607.22770
作者: Siyuan Du,Mengxi Chen,Xinyang Jiang,Zilong Wang,Jiangchao Yao,Dongsheng Li,Ya Zhang,Lili Qiu,Yanfeng Wang
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Although artificial intelligence (AI) has shown promising performance in several medical tasks, accurate dementia etiology diagnosis with AI remains challenging due to complex overlapping symptoms among diseases. Scaling up the dataset size by combining the cross-center samples may bring a gain in the pursuit of performance, while the inherent data heterogeneity across centers or populations induces the conflict. Conventional multi-task learning paradigms offer a promising framework; however, they fail to consider critical meta information (e.g., site-specific acquisition and modality availability) to combat the heterogeneity. To address this challenge, we propose a Collaborative Meta Knowledge Enhancement (COME) framework for dementia etiology diagnosis, which injects multi-center acquisition semantics, source identifiers, and modality indicators as heterogeneity-aware embeddings into a unified Transformer architecture for scale-up training, enabling explicit modeling of heterogeneity. Besides, a trust-region constrained optimization scheme is designed to regularize the model from spurious correlations during training through a reference model. Across seven independent cohorts, our method achieves state-of-the-art in-domain performance with a mean macro-averaged AUC of 85.62% and a 4.29-point gain over the strongest baseline, while maintaining superior out-of-domain generalization under both cross-center and cross-sequence evaluations. Extensive validation also confirms the alignment between model predictions and established biomarkers (amyloid, tau) and clinical severity, highlighting the potential of COME to enable robust and interpretable dementia diagnostics in real-world settings.

[CV-165] Real-time Reconstruction of Human Visual Perception from fMRI

链接: https://arxiv.org/abs/2607.22753
作者: Rishab S. Iyer,Jiaxin Cindy Tu,Cesar Kadir Torrico Villanueva,Anish Mahishi,Ross P. Kempner,Jacob S. Prince,Ernest W. Lo,Akash Bhowmick,Hritik Arasu,Amaar Chughtai,Elizabeth A. McDevitt,Paul S. Scotti,Kenneth A. Norman
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)
备注:

点击查看摘要

Abstract:Real-time closed-loop neurofeedback based on functional magnetic resonance imaging (fMRI) has led to important scientific and clinical advances. However, the sophistication of the analysis methods used in real-time fMRI lags behind the state-of-the-art in fMRI decoding, largely due to computational factors: Most advanced decoding pipelines do not fit within the envelope of real-time processing, where the analysis needs to be conducted in a matter of seconds and without leveraging data acquired later in the session. Here, we present a real-time compatible adaptation of a computationally intensive state-of-the-art pipeline for reconstructing perceived natural images (MindEye2), and we demonstrate that reliable fine-grained decoding is still achievable in this setting. Using RT-Cloud, an open-source, scalable cloud-based platform, we performed a real-time scan where we decoded single-trial visual perception within seconds after an image was shown to the participant. Finally, we use simulated analyses to document the factors driving changes in performance from offline to real-time analysis. This work serves as a proof-of-concept that it is feasible to deploy these powerful fMRI decoding pipelines in real-time analysis, paving the way for their use in brain-computer interfaces for scientific discovery and clinical treatment.

[CV-166] Beyond Error-vs-Discard Characteristic: Toward Stable and Reliable Evaluation for Face Image Quality Assessment

链接: https://arxiv.org/abs/2607.22752
作者: Bhavesh Wani,Žiga Babnik,Vitomir Štruc,Philipp Terhörst
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Face Image Quality Assessment (FIQA) aims to estimate the utility of facial images for reliable recognition. The evaluation of FIQA methods is predominantly based on the Error-versus-Discard Characteristic (EDC), which evaluates performance by progressively discarding low-quality samples and measuring recognition error on the retained subset. In this work, we demonstrate that the widely used EDC protocol has fundamental limitations: Test-Set Divergence and Threshold Drift, which together limit the reliability and comparability of FIQA methods. To address this, we propose discard-based EDC variants and a rank-based Rank Consistency Evaluation (RCE) metric that operates on the entire test set without discarding samples, using a fixed decision threshold. Extensive experiments on five datasets, four face recognition models, and 15 state-of-the-art FIQA methods demonstrate both the limitations of EDC and the effectiveness of the proposed approaches in enabling a more reliable and comparable evaluation. Despite evaluated on face images only, the limitations arise from the EDC protocol rather than the biometric modality, suggesting a broader applicability to biometric quality assessment in general.

[CV-167] Post-Operative Glioma Segmentation via Loss Stabilization Normalization and Subspace Attention ICANN2026

链接: https://arxiv.org/abs/2607.22749
作者: Alexandru Crişan,Diana Borza
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted for presentation at the International Conference on Artificial Neural Networks (ICANN 2026)

点击查看摘要

Abstract:Tracking residual tumor after surgery is essential for catching recurrence early, but automating post-operative glioma segmentation remains a difficult task. Although transformer-based architectures, such as SwinUNETR, achieved impressive results, few studies test how well they generalize across clinical protocols. In this paper, we conduct an ablation study on the MU-GLIOMA-POST and UCSF-ALPTDG datasets and show that the standard Generalized Dice Loss (GDL) is unstable under domain shift: the Whole Lesion (WL) Dice drops from 0.88 on the internal validation set to 0.73 on the external UCSF test set. To address this, we pair brain-masked percentile normalization with voxel-level contrastive learning. We also propose a Subspace-Aware Class Attention (SACA) module that re-calibrates the bottleneck features and raises Enhancing Tumor (ET) sensitivity by 8% (9.1% relative improvement) on internal validation. Ensembling these refinements with nnU-Net brings every stable configuration to a WL Dice of 0.94, and the SACA variant ensemble achieves the best boundary error (HD95) of 2.92 mm on MU-GLIOMA-POST.

[CV-168] Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge

链接: https://arxiv.org/abs/2607.22746
作者: Hongruixuan Chen,He Huang,Haifeng Wang,Jian Song,Junjue Wang,Weihao Xuan,Hamish Mitchell,Jiepan Li,Wei He,Liangpei Zhang,Zijie Wang,Chen Zhong,Jiazhen Zhao,Lei Hu,Ting Hu,Hongyan Zhang,Gregory Angelides,Miriam Cha,Clifford Broni-Bediako,Junshi Xia,Taylor Perron,Naoto Yokoya
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV)
备注:

点击查看摘要

Abstract:Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed. Post-event optical imagery, however, may be unavailable because of cloud, smoke, or darkness. The Bright Challenge evaluated all-weather building damage mapping from a submeter-resolution pre-event optical image and a post-event SAR image. Participants were required to detect and delineate each building and assign exactly one of three mutually exclusive damage labels. The challenge extended the globally distributed \textscBright dataset with instance-level annotations for about 291,000 buildings across 16 disaster events spanning seven disaster types. The final phase was evaluated exclusively on two 2025 events absent from training: a wildfire event in California and a hurricane in Jamaica. A total of 157 participants made 1,289 submissions, and 46 teams entered the final phase. The two winning solutions achieved test mAPs of 0.182 and 0.181, approximately 8.7 times the public baseline of 0.021, but remained far below the best in-domain holdout score of 0.513. Across teams ranked in both phases, performance declined sharply and the rank order changed substantially. The two leading solutions independently favored modality-specific encoding, staged or late optical–SAR fusion, and an optical-dominant separation of building localization from damage recognition. The winning method additionally used scene-aware threshold adjustment and pseudo-label adaptation. These results identify cross-event generalization and stable severity discrimination as the principal remaining challenges. All data, annotations, baseline code, and winning solutions are publicly available at this https URL.

[CV-169] AI-generated Images Challenge Visual Trust in High-risk Scenarios

链接: https://arxiv.org/abs/2607.22745
作者: Yi-Zhi Wang,Yichen Xiao,Linan Yue,Weibo Gao,Yichao Du,Pengfei Fang,Shimin Di,Min-Ling Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation. Yet existing detection benchmarks rarely examine synthetic images in public- and individual-safety contexts, where misleading visual content may carry substantial risks. Here we introduce SafeIMG, a safety-oriented benchmark spanning 12 public- and individual-safety scenarios generated using GPT Image 2. Unlike benchmarks centred on generic imagery and image-level labels, SafeIMG evaluates not only whether detectors recognise synthetic images, but also whether their decisions reflect human-identified anomalies. To this end, SafeIMG provides human annotations that localise suspicious regions and explain local artefacts and higher-level commonsense or physical inconsistencies. We evaluate specialized synthetic-image detectors and vision-language models (VLMs), and find that neither provides reliable detection. The strongest VLM identifies only 49.5% of generated images, whereas the best specialised detector identifies 33.1%, compared with 81.7% accuracy for human evaluators. Model explanations cover only 29.8% of human-annotated anomalies and predominantly capture local defects in text, faces and hands. Their coverage falls to 15.0% for commonsense conflicts and 12.0% for physical inconsistencies, while detection performance deteriorates further after dissemination-induced image degradation. These findings show that current detectors lack the accuracy, explanatory alignment and robustness needed to evaluate AI-generated images reliably across public- and individual-safety settings.

[CV-170] QFedPolyp: A Communication- and Inference-Efficient Federated Learning Framework for Polyp Segmentation

链接: https://arxiv.org/abs/2607.22743
作者: Madan Baduwal,Priyanka Paudel
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Background and Objective: Automatic polyp segmentation supports computer-aided diagnosis and early colorectal cancer detec- tion. Centralized deep learning requires hospitals to share sensitive medical data, while federated learning preserves privacy but introduces high communication costs through repeated transmission of full-precision model parameters. We propose QFedPolyp, a communication- and inference-efficient federated learning framework for collaborative polyp segmentation. Methods: QFedPolyp combines quantization-aware training with low-precision model communication. Each hospital locally trains a lightweight U-Net on private data while simulating quantization during training. Clients transmit quantized model parameters to a central server, where they are reconstructed and aggregated using Federated Averaging. Evaluation is performed on Kvasir-SEG, CVC-ClinicVideoDB, PolypGen, and BKAI-IGH NeoPolyp. Results: Full-precision federated training achieves Dice scores of 0.910 on Kvasir-SEG and 0.930 on CVC-ClinicVideoDB. Uni- form 8-bit communication reduces transmission cost by approximately 4 times while preserving competitive segmentation accuracy. Quantized models also achieve up to 1.5 times faster inference than full-precision models. Conclusions: QFedPolyp enables privacy-preserving collaborative polyp segmentation with reduced communication overhead and faster inference. The resulting lightweight models are suitable for real-time clinical deployment. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.22743 [cs.LG] (or arXiv:2607.22743v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.22743 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Madan Baduwal [view email] [v1] Thu, 23 Jul 2026 01:50:54 UTC (12,763 KB)

[CV-171] A Diagnostic Gap Framework for Evaluating Reconstruction Fidelity in Weakly Supervised Mammography

链接: https://arxiv.org/abs/2607.22740
作者: Vinceline Bertrand,Ionut Cardei
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Weakly supervised pipelines for medical imaging have become increasingly popular over the years. These systems often include multiple stages and components, such as reconstruction, generation, and localization, yet standard evaluation metrics provide limited insight into whether clinically relevant information is preserved across each stage. We present the diagnostic gap framework, a practical evaluation tool that measures decision preservation and explanation preservation as a function of measured reconstruction fidelity. To isolate the effect of reconstruction from localization, we evaluate on curated lesion ROI crops using a fidelity ladder of three class-conditional reconstructors—VQ-VAE-GAN, VAE-GAN, and diffusion (SDEdit)—spanning a twenty-fold range in perceptual distance (LPIPS 0.029–0.584). At autoencoder fidelity, both decision and explanation are preserved: AUC changes remain within \pm 0.005 and attribution similarity (HiResCAM, Grad-CAM++) stays high. At diffusion fidelity, both collapse: pooled AUC drops by 0.253 and mass-pathology AUC falls below chance. The diagnostic gap is thus a measurable function of reconstruction fidelity rather than an intrinsic cost of reconstruction, and the framework provides an architecture-agnostic instrument for identifying when and where multi-stage pipelines lose diagnostic signal.

[CV-172] Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features

链接: https://arxiv.org/abs/2607.22739
作者: Dzmitry Malyshau
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 13 pages, 2 figures, made with heavy assistance from LLM

点击查看摘要

Abstract:We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcement learning or explicit memory. Cortex is a compact Quake policy with 10.98 million trainable parameters in a six-layer transformer over a frozen DINOv3 encoder. It is trained on the Quake subset of the public Pixels2Play corpus: 6,849 recordings (about 474.7 hours), represented as 17.09 million cached decision frames with keyboard and mouse actions. One sampled training epoch uses 517,048 four-frame windows and takes 3.3 minutes of policy-head optimization on one RTX 5080, excluding one-time feature extraction. We evaluate two independent batches of 20 stochastic, 120-second episodes on Quake E1M1. Cortex does not complete the level, but every episode reaches the opening door, button room, and gate descent; 19 of 20 episodes in each batch record at least one kill. Under the same time-controlled harness, released P2P-150M and NitroGen checkpoints remain shallower in five matched-duration episodes each. These comparisons are limited by small reference samples and different native interfaces. Ablations show that denser visual tokens improve combat and survival, while longer optimization and naive action history improve offline metrics without consistently improving play. The remaining failures are consistent with covariate shift and motivate targeted corrective data. We release the policy implementation, checkpoint, and a representative rollout.

[CV-173] Nova3D: Code-Native Generation of Programmable 3D Assets

链接: https://arxiv.org/abs/2607.22738
作者: Nimra Noor,Muhammad Bilal,Abdullah Hussain,Hassan Baig
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Current 3D generative models mostly produce a final surface: a visually strong but largely opaque mesh. Interactive 3D worlds need more than a surface. They need named parts, an assembly hierarchy, measurable constraints, local edit handles, and joints for articulation. We present Nova3D, a system that generates 3D assets as executable Blender source code; the compiled mesh, a binary glTF (GLB), is treated as the artifact, not the asset. Because the output is a program, semantic handles exist at generation time rather than being recovered afterward by segmentation or rigging. We evaluate on Nova3D-Bench, a frozen, spec-grounded benchmark of 54 items across six domains and three difficulty levels with text and image inputs, against eleven baselines in four families (mesh-native, part-structured, code-native, and CAD) plus a same-LLM ablation. Nova3D produces an executable program and a valid artifact for 54/54 items. Every asset exposes named parts organized in a parent-child assembly tree; no mesh-native, CAD, or segmentation baseline exposes either. It satisfies 51/52 prompt-stated numeric and count constraints (best baseline: 11/52), passes 14/18 blinded local edits with locality preserved in 18/18, and articulates 59 joints across 12 assets at 98.3% geometric validity, where every baseline exposes zero native joints. Its geometry is competitive: it wins the structured domains in a pairwise shape-quality tournament and is second only to the strongest mesh-native model, while conceding texture realism to baked-PBR systems. The central result is representational: code-native generation turns a generated 3D object from an opaque surface into a programmable asset that downstream systems can inspect, measure, edit, and animate.

[CV-174] Benchmarking the Domain Gap: Model Selection Instability Under Domain Shift in Video Capsule Endoscopy

链接: https://arxiv.org/abs/2607.22736
作者: Dan Hanson,Debesh Jha
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages

点击查看摘要

Abstract:Video capsule endoscopy (VCE) classification is typically evaluated within a single dataset, yet clinical deployment demands robustness across acquisition sources, labeling policies, and patient populations. We examine this gap using Kvasir-Capsule, Capsule Vision 2024 (CV2024), and a shared-label subset of Galar. We fine-tune a suite of general-domain pretrained backbones on the official Kvasir-Capsule folds under a standardized protocol and evaluate the same checkpoints on two non-source targets within a documented shared-label decision space. We find that the predictive value of in-domain ranking is target-dependent: Kvasir-Capsule ranking aligns more closely with Galar than with CV2024, while the two non-source targets agree only weakly. Consequently, the strongest in-domain backbone leads on one target yet falls to mid-pack on the other, and no single evaluation target reliably predicts the others. A second CV2024-trained configuration set reproduces this target-dependent instability. We conclude that capsule endoscopy model selection should report cross-target ranking stability rather than peak single-dataset performance.

[CV-175] Fast Fourier Convolutional GAN for 30 m Clear-Sky Land Surface Temperature Gap-Free Reconstruction

链接: https://arxiv.org/abs/2607.22734
作者: Marwa Alfouly,Smajil Halilovic,Nils Bochow,Thomas Hamacher,Niklas Boers,Konrad Schindler
类目: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV); Geophysics (physics.geo-ph)
备注: 35 pages, 9 figures, Journal

点击查看摘要

Abstract:Satellite-derived Land Surface Temperature (LST) provides spatially comprehensive data that ground stations cannot match. However, its utility is frequently limited by severe data gaps due to the presence of clouds. As LST is essential for understanding land-atmosphere interactions, numerous methods have been proposed to address this challenge. Yet, the development of a scalable and adaptable pipeline for generating gap-free LST datasets and reconstructing cloud-contaminated pixels remains challenging. Moreover, the reconstruction of extensive missing regions in fine-spatial-resolution observations is particularly difficult. To address this challenge, we propose a Multimodal Fast Fourier Convolutional GAN for reconstructing cloud-contaminated pixels in fine-resolution (30 m) Landsat imagery to generate gap-free clear-sky LST products. The method leverages Fast Fourier Convolution to enable a global receptive field across the image, and is guided by a stack of data consisting of satellite observations and Synthetic Aperture Radar (SAR) data. Across all LST quantiles, the interquartile range of scene-averaged RMSE (computed over reconstructed pixels) is consistently between 0.8 K and 1.8 K. The proposed approach enables the recovery of extensive missing regions, including scenes with more than 70% cloud-induced gaps, while relying on auxiliary data that are readily available at a near-global scale.

[CV-176] Generative Augmentation for EEG Motor Imagery Classification: A Class-Conditional VAE with Cycle-Consistent Decoder Refinement

链接: https://arxiv.org/abs/2607.22733
作者: Matei Moldoveanu,Alain Sirois,Claire Ben Ali,Fabien Lotte,Florian Yger
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:We investigate whether a generative model can supply useful synthetic motor-imagery (MI) electroencephalography (EEG) trials that improve the accuracy of independent downstream classifiers. We train a class-conditional variational autoencoder (CVAE) with an integrated latent classifier on the Zhou motor-imagery dataset, using the learned per-class prior as a generator: sampling the prior for a given label and decoding it into a synthetic, label-consistent signal. A constraint on the covariance matrix of the generated data encourages preservation of covariance structure, and the model is trained with a schedule that alternates ordinary VAE training with a decoder-focused phase that sharpens the generative pathway used for augmentation. We measure the effect of adding synthetic trials to the training set under two evaluation protocols – within-user (pooled 60/20/20 split across subjects) and cross-user (leave-one-subject-out, LOSO) – across four representative EEG classification pipelines: Common Spatial Patterns with Linear Discriminant Analysis (CSP+LDA), tangent-space features with a Support Vector Machine (TGSP+SVM), Minimum Distance to Riemannian Mean (MDM), and a neural network based on EEGNetv4 (henceforth EEGNet). Results are aggregated across independent augmentation draws, random seeds (within-user), or leave-one-subject-out folds (cross-user), with uncertainty reported as 95% confidence intervals (Student’s t -distribution) computed over per-seed/per-fold averages. We find that synthetic EEG from the CVAE is most credible as a source of class-structured, covariance-like data rather than as a substitute for real raw EEG: it can raise the point estimate for MDM, but the broader augmentation claim remains conservative – observed gains are small and classifier-dependent.

[CV-177] Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment

链接: https://arxiv.org/abs/2607.22731
作者: Ponleur Veng(CADT, M-PSI),Dominique Vaufreydaz(LIG, M-PSI),Phutphalla Kong(CADT)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Multi-Camera People Tracking (MCPT) traditionally relies on precise intrinsic and extrinsic camera calibration to project 2D detections into a unified 3D world coordinate this http URL, manual calibration constitutes a major bottleneck in large-scale dataset generation from unconstrained video archives. This work proposes a unified calibration-free 3D MCPT framework that infers geometric structure directly from visual data using deep foundation models. The system integrates anchor-free detection (YOLOX), robust tracking (BoT-SORT), omni-scale appearance embedding (OsNet), pose estimation (HRNet via MMPose), and transformer-based geometric reconstruction using the Visual Geometry Grounded Transformer (VGGT). A pose-guided 3D lifting strategy projects head keypoints onto a reconstructed manifold, eliminating dependence on ground-plane homography. Global identity association is formulated as hierarchical agglomerative clustering under a joint appearance-geometry cost with strict velocity gating. Evaluation on the AI City Challenge 2024 demonstrates a HOTA score of 53.13% without access to ground-truth calibration matrices, establishing a strong baseline for purely vision-based 3D tracking.

[CV-178] CrossSpine: Multi-scale Cross-sequence Attention with Anatomical Priors for Automated Pfirrmann Grading

链接: https://arxiv.org/abs/2607.22728
作者: Hai Son Nguyen,Duong Ngoc Vu,Trong-Nghia Nguyen,Bien Tran Van,Van-Dem Pham,Trang Mai Xuan,Huan Vu,Thien Van Luong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Automated grading of Lumbar Disc Degeneration is essential for the objective quantification of structural changes associated with low back pain. Observing that baseline models underperformed on our data, we propose a framework designed to overcome these limitations. First, we present the Cross-sequence Attention Spine (CrossSpine) framework, a novel architecture that employs a cross-sequence attention mechanism to adaptively fuse features from different MRI sequences at multiple spa- tial scales. Second, we contribute a meticulously curated dataset aimed at automated Pfirrmann grading. Finally, we introduce an IVD-aware classification technique that integrates anatomical disc-level information, enabling the model to learn level-specific degeneration priors. Our experi- ments demonstrate the superiority of this approach: CrossSpine achieved a relative improvement exceeding 125% in the Macro F1 score, while boosting the Mean AUPRC by 99% and the Mean AUROC by 36% com- pared to the baseline.

[CV-179] rustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation

链接: https://arxiv.org/abs/2607.22727
作者: Pranav Kaliaperumal,Manisha Kaliaperumal
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning. We present a reproducible framework for evaluating uncertainty-aware segmentation under con- trolled clinical degradation. Our experiments use a synthetic multimodal brain tumor MRI cohort generated with a biophysical phantom simulator that follows the BraTS protocol. We train U-Net and Attention U-Net baselines for multi-class tumor sub-region segmentation and augment both models with Monte Carlo dropout to estimate per-voxel uncertainty. Across eight clinically motivated corruption types at five severity levels, we measure segmentation accuracy, calibration, failure detection, and selective prediction coverage. On clean data, Attention U-Net achieves a whole-tumor Dice of 0.990; under severe Gaussian noise, its performance falls to 0.089. Predictive uncertainty rises with degradation and tracks segmentation error (Pearson r = 0.53 under severity-3 Gaussian noise), allowing us to flag failures with an AUROC of 0.843. These results argue for uncertainty-aware inference as a practical safety layer in physician-in-the-loop radiology workflows. We release the code, trained models, and evaluation protocol to support direct reproduction.

[CV-180] PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models ACM-MM2026

链接: https://arxiv.org/abs/2607.22726
作者: Zihan Song,Shuo Ye,Bo Zhao,Ruixin Zhang,Jiayu Zhang,Shouhong Ding,Zitong Yu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ACM MM 2026

点击查看摘要

Abstract:Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindrance to efficient reasoning. This paper introduces a training-free \mathbfP ersistence-Aware \mathbfC ompression and \mathbfA ggregation (PCA) method designed to preserve high-fidelity raw visual information before the encoding stage. PCA can be built on arbitrary VLLMs and consists of two modules: 1) A Dynamic Downsampling (DD) module that adaptively removes redundant frames by analyzing frame-wise similarity. 2) A Persistence-Aware Motion Enhancement (PAME) module that enriches each selected keyframe by aggregating the temporal context of its neighbors, ensuring that essential information is preserved even after aggressive frame reduction. Our approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model. Extensive experiments demonstrate that PCA consistently outperforms existing state-of-the-art approaches in both efficiency and accuracy, achieving a speedup of 1.8 \times to 2.5 \times compared to the baseline VLLM. The code is open-sourced at this https URL.

[CV-181] Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification

链接: https://arxiv.org/abs/2607.22725
作者: Mohamed Abdallah Salem,Nourhan Zein Diab
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Copyright 2026 IEEE. This is the author’s version of the work that has been Accepted for publication in the Proceedings of the 2026 IEEE 2025 Intelligent Methods, Systems, and Applications (IMSA). Final published version will be available on IEEE Xplore

点击查看摘要

Abstract:Data augmentation is routinely used to improve generalization in image classification, but the assumptions underlying standard policies are poorly matched to coherent imaging. Laser speckle patterns are not generic textures; they arise from coherent interference, and their discriminative content is carried by structured stochastic spatial and frequency statistics. This study examines how controlled augmentation perturbations influence speckle-based material classification on the SensiCut dataset. We train ResNet18 and EfficientNet-B0 under a parametric augmentation framework comprising rotation, Gaussian blur, independent Gaussian noise, spatially correlated speckle-aware noise, intensity jitter, and spatial masking, and evaluate test performance using macro F1-score averaged over three random seeds. Separate ordinary least squares models link augmentation parameters to performance for each architecture. Across both models, Gaussian blur exerts a strong negative effect (p 0.001), indicating that low-pass filtering suppresses high-frequency structure that is informative for material discrimination. Independent pixel-wise noise is likewise harmful (p = 0.003 for EfficientNet-B0 and p = 0.001 for ResNet18), consistent with disruption of local spatial coherence. In contrast, spatially correlated perturbations yield significant positive coefficients (p = 0.004 for EfficientNet-B0 and p = 0.001 for ResNet18), showing that variability can improve robustness when it preserves speckle organization. The fitted models explain a substantial fraction of performance variation (R2 = 0.796 for EfficientNet-B0 and R2 = 0.879 for ResNet18). These results show that, in laser speckle imaging, augmentation effectiveness is determined primarily by structural preservation rather than perturbation magnitude. The findings motivate physics-aware augmentation design for coherent optical sensing.

[CV-182] Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models

链接: https://arxiv.org/abs/2607.22723
作者: Huafu Li,Guo Chen,Jia Xia,Lei Wang,Wei Du,Yun Yao,Weijun Peng,Liming Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 14 pages, 3 figures

点击查看摘要

Abstract:Visual information extraction (VIE) from visually rich documents remains challenging due to high layout variability and real-world impairments. Existing methods typically rely on sequential OCR pipelines or end-to-end models requiring extensive labeled data and layout-specific training, limiting their this http URL propose a classification-guided large vision-language model (LVLM) framework for multi-type VIE that achieves high accuracy with minimal supervision. The approach decouples document-type classification from content extraction and employs in-context learning (ICL)-based dynamic prompt engineering to inject task-specific knowledge, enabling robust zero-shot inference across diverse layouts. From a theoretical perspective, the proposed method can be viewed as a form of conditional computation that reduces task uncertainty and improves information efficiency during prompt-based inference. Evaluated on a real-world bidding dataset with 16 certificate types, our zero-shot method (based on Qwen2.5-VL-7B) outperforms a strong supervised baseline by 18.35 percentage points in F1-score (86.43% vs. 68.08%) and 0.23 in normalized edit distance (0.90 vs. 0.67). Optional domain-specific fine-tuning further improves performance to 93.65% F1 and 0.93 NED, demonstrating superior robustness against seals, watermarks, and low contrast. The framework offers an efficient, scalable solution for complex document understanding in office automation. Code is available at this https URL, and fine-tuned models at this https URL.

[CV-183] A New Kind of Adversarial Example: Measuring the Human-Model Gap and Its Relationship to OOD Detection

链接: https://arxiv.org/abs/2607.22722
作者: Ali Borji
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation that causes the model to keep its original, correct prediction, even though a human would no longer recognize the image. Prior work showed such examples can be generated at scale but left three questions untested: whether humans really perform worse than the model, whether standard out-of-distribution (OOD) detection and calibration tools catch it, and whether existing defenses mitigate it. We answer all three on MNIST, CIFAR-10, and ImageNet. (i) An independent recognizer proxy drops to ~49% on CIFAR-10 while the model stays at 100% – a gap a small human pilot (N=5) corroborates directly and that is not explained by signal loss (a matched-magnitude Gaussian control degrades recognizability faster); a CLIP zero-shot proxy confirms the gap at ImageNet scale too. (ii) Confidence- and energy-based OOD detectors and calibration are structurally blind (0% detection, ECE ~= 0), while a feature-space Mahalanobis detector flags 100% – but is evaded by an adaptive attacker at no cost to success. (iii) No classical defense, including adversarial training (45% robust accuracy), reduces attack success (correlation with large-epsilon_l resistance r ~= 0). A mechanistic analysis further shows the attack destroys low-level texture far faster than edge/shape structure.

[CV-184] An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia

链接: https://arxiv.org/abs/2607.22721
作者: Nassira Ait Mehdi,Milissa Temmam,Slimane Larabi
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning. For patients diagnosed with schizophrenia, these tasks are crucial for addressing severe cognitive deficits. However, evaluating the correctness of these physical actions generally relies on manual observation by clinicians, which introduces subjectivity and limits the scalability of therapeutic interventions. In this paper, we propose an automated framework based on Vision-Language Models for action verification in cognitive remediation tasks tailored for schizophrenia rehabilitation. The proposed system relies on a camera-monitored tabletop environment composed of structured miniature scenes including roads, a roundabout, a park, and toy vehicles. Patients receive audio instructions describing goal-oriented spatial actions to perform by manipulating a toy vehicle. These interactive physical activities are specifically designed to stimulate targeted cognitive functions, such as sustained attention, motor coordination, spatial navigation, and cognitive flexibility. To verify the correctness of the performed actions without requiring continuous clinical oversight, the system analyzes the video feed tracking the patient’s hand and toy movements. A fine-tuned Vision-Language Model interprets the recorded video sequences and generates semantic descriptions of the observed activities, enabling high-level verification of the executed actions with respect to the initial textual instructions. A dedicated dataset of 4634 tabletop cognitive remediation video scenarios was collected to evaluate the proposed approach. Experimental results demonstrate that our specialized framework effectively bridges low-level physical telemetry with high-level clinical feedback, presenting a scalable and objective solution for advanced cognitive rehabilitation.

[CV-185] γ-Bridge: A Look-Parametric Diffusion Bridge

链接: https://arxiv.org/abs/2607.22719
作者: Xuran Hu,Yujie Zhu,Tengxi Wang,Jilong Li,Wufan Zhao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 13 pages, 9 figures

点击查看摘要

Abstract:Multiplicative Gamma noise is a signal-dependent degradation in coherent imaging; synthetic aperture radar (SAR) despeckling is its most prominent real-world instance. Existing diffusion denoisers parameterize their forward process by abstract signal-to-noise schedules rather than by the physical look number L , so different deployment scenarios typically require separately trained models, and transfer from synthetic Gamma training to real SAR remains challenging without clean ground truth. We introduce \gamma -Bridge, a look-parametric bridge whose schedule L(t) connects the noisy observation at L_obs to the clean limit through exact multiplicative Gamma marginals. Its closed-form Gamma–Lévy reverse posterior admits both stochastic and deterministic processes, while observation conditioning and a two-step consistency loss stabilize multi-step inference in the low-SNR single-look regime. Because bridge time directly represents L , one conditioned network can smart-start from any admissible input look and stop at a target look number. These two orthogonal controls enable zero-shot restoration over the full admissible grid after training only at L_obs = 1 on natural images with synthetic Gamma corruption. Combined with a homogeneous-patch look estimator, \gamma -Bridge processes data from six spaceborne and airborne SAR sensors without sensor-specific fine-tuning, achieving leading results on standard synthetic benchmarks while providing physically interpretable input and output controls absent from prior denoisers. Codes are released \hrefthis https URLhere.

[CV-186] DAMamba-UNet3D: A Parameter-Efficient Mamba State Space U-Net with Dynamic Adaptive Scan for 3D Medical Image Segmentation

链接: https://arxiv.org/abs/2607.22718
作者: Mohammad Arafat Hussain,Ellen Grant,Yangming Ou
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We propose parameter-efficient SSM-based U-Net architectures for 3D medical image segmentation. Convolutional U-Nets afford O(n) local mixing per layer but lack explicit global context; transformers provide global reasoning at O(n^2) cost in sequence length n . State-space models (SSMs), such as Mamba, offer O(n) global propagation per block. Yet, existing medical SSM segmenters rely on fixed scan patterns and large parameter budgets. Dynamic Adaptive Scan (DAS), which learns data-dependent reordering before selective scan, has not been applied to medical imaging or extended to 3D volumes. We propose DAMamba-UNet3D, a hybrid encoder-decoder that integrates tri-plane 3D-DAS blocks at encoder stages E2-E4 while retaining convolutions elsewhere (~5.3M parameters). On BraTS 2020 five-fold cross-validation, DAMamba-UNet3D achieves mean Dice 0.815+/-0.013 (full-volume per-case evaluation) at ~13x lower parameter cost than SegMamba (0.824+-0.014, ~70M). At comparable scale, DAMamba-L (~70M), a wide DAS-native variant with encoder-only DAMamba and a convolutional bottleneck, reaches 0.829+-0.012, surpassing retrained SegMamba by 0.5pt. Component ablations show that encoder-only DAS placement is critical as bottleneck and decoder SSM blocks lower Dice. Together, the results suggest that learned tri-plane DAS in a hybrid U-Net is competitive with, and under our large-scale design may improve upon, SegMamba’s fixed Tri-orientated Mamba (ToM) scanning on BraTS 2020. Code: this https URL.

[CV-187] OM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians

链接: https://arxiv.org/abs/2607.22717
作者: Marek Lisowski,Łukasz Smoliński,Kornel Howil,Piotr Biliński,Marcin Mazur,Przemysław Spurek
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:While Implicit Neural Representations (INRs) and dynamic 3D Gaussian Splatting (3DGS) achieve impressive results in video processing, they often fall short of producing representations that are easily editable. Recent methods address this by introducing complex spatial deformations or folded distributions, which constrain optimization and reduce flexibility for downstream editing. In this paper, we introduce TOM-GS, an editable video representation that forgoes complex deformations in favor of regular 3D Gaussians equipped with a continuous temporal opacity formulation. By assigning a learnable temporal mean and scale to the opacity of each Gaussian, our model enables static 3D spatial components to fade smoothly in and out of the scene. Grounded by robust, off-the-shelf pose estimation, our approach maintains a static spatial geometry that naturally supports a wide range of manual and physics-based edits. TOM-GS outperforms prior editable video representations in visual fidelity, while its reliance on standard 3D Gaussians ensures seamless compatibility with established 3D editing tools.

[CV-188] Visual Token Compression Enhances Robustness of MLLM s

链接: https://arxiv.org/abs/2607.22716
作者: Shishen Gu,Jiequan Cui,Wenbo Hu,Zenglin Shi,Zhenzhen Hu,Richang Hong
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 20 pages, 16 figures. Accepted at ACM Multimedia 2026. Code: this https URL

点击查看摘要

Abstract:In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations. Given that vision and language modalities cannot be perfectly aligned, the misaligned visual tokens might act as out-of-distribution (OOD) inputs, leading to unpredictable outputs and introducing potential vulnerabilities. Building on this insight, we aim to enhance model robustness against jailbreaks and hallucinations by reducing OOD visual tokens at robust-pruning layers, while also reducing inference cost as a side benefit. Specifically, we measure the distance between each visual token and the language feature space. Then, visual tokens with large distances are identified as OOD tokens, which can be iteratively pruned. To demonstrate the effectiveness of our method, we evaluate it on seven diverse popular benchmarks. Notably, our method yields an average improvement of 13.29% in defending jailbreak attacks, consistently achieves competitive performance in mitigating hallucinations, and maintains strong results on general datasets like MME.

[CV-189] Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework

链接: https://arxiv.org/abs/2607.22715
作者: Lewen Mi,Manyi Li,Yuling Sun,Yufan Zhang,Yuxin Shi,Yulong Bian,Xiangxian Li,Juan Liu
类目: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:The rapid growth of Artificial Intelligence-generated content (AIGC) is reshaping video production and circulation, exposing children to an increasing volume of AIGC videos. Unlike traditionally produced videos, AIGC videos often exhibit greater uncertainty in visual details, narrative coherence, and content expression, which may introduce developmentally inappropriate risks for children. However, existing video safety research is largely designed for general violation detection from an adult perspective and remains insufficient for identifying the fine-grained, implicit, and context-dependent risks that children may encounter when viewing AIGC videos. To address this gap, we study child-oriented AIGC video reviewing, making three contributions. First, we construct CAVSR, a benchmark of 605 real-world videos collected from multiple platforms, and develop a hierarchical risk taxonomy comprising 6 top-level categories and 26 fine-grained labels to support systematic evaluation of children’s viewing risks. Second, we propose QVRS-E, a knowledge- and experience-augmented video reviewing framework that combines multi-agent collaboration with expert and experiential knowledge to support targeted evidence acquisition and fact-grounded reviewing decisions. Third, extensive experiments demonstrate that our method significantly enhances the reviewing of child-related risks integrated with vision-language models, and yields more robust review reports.

[CV-190] Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems

链接: https://arxiv.org/abs/2607.22714
作者: Sai Sidharth D
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive platforms impose severe constraints on compute, memory, and power. This paper presents an optimized semantic segmentation architecture derived from the RetinaNet detection framework, adapted for dense pixel-wise prediction and tailored for deployment on resource-constrained embedded hardware. The proposed architecture, termed Opt-RetinaSeg, replaces the standard ResNet-50 backbone with a hybrid lightweight feature extractor, restructures the Feature Pyramid Network (FPN) to reduce redundant multi-scale computation, and introduces a compact segmentation head guided by focal-loss-inspired class balancing to address the severe foreground-background imbalance common in road scenes. We further apply a three-stage optimization pipeline consisting of structured channel pruning, post-training INT8 quantization, and knowledge distillation from a high-capacity teacher network. Evaluated on the Cityscapes and BDD100K datasets and deployed on an NVIDIA Jetson Xavier NX and a Qualcomm QCS610 automotive SoC, the proposed model achieves 73.9% mIoU at 70.4 FPS, representing a 7.4x inference speedup and a 4x reduction in model size relative to the ResNet-50 baseline, with less than 3% accuracy degradation. These results indicate that RetinaNet-derived architectures, when systematically optimized, are viable candidates for real-time semantic segmentation in embedded automotive perception pipelines

[CV-191] scMIR: a vision-language foundation model for single-cell light microscopy image representation

链接: https://arxiv.org/abs/2607.22712
作者: Yifan Shang(1 and 2),Jiahui Tan(2),Xiangxiang Zeng(2),Renjie Zhou(1) ((1) Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong, China, (2) College of Computer Science and Electronic Engineering, Hunan University, Changsha, China)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Optics (physics.optics)
备注:

点击查看摘要

Abstract:Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges to high-throughput automated analysis. Existing representation learning methods mostly rely on task-oriented modeling, which is limited by specific datasets and predefined tasks, making them difficult to generalize across different cell types and microscopy modalities, and experimental conditions. Although general-purpose methods have improved the generalization ability of image representation in recent years, their limited utilization of experimental background and biological context information still poses challenges in complex phenotypic analysis. Here, we propose scMIR, a vision-language foundation model for single-cell light microscopy image representation. By synergistically combining self-supervised image reconstruction with text-guided cross-modal alignment, scMIR can simultaneously encode morphological and biological semantic information in a unified representation space. scMIR is pre-trained on 207,957 image-text pairs, covering various cell types, microscopy modalities, and perturbation conditions. scMIR outperforms existing general models and task-oriented methods as systematically evaluated on various complex tasks using 16 benchmark datasets, including cell classification, clustering, phenotype inference, and batch effect correction tasks. Furthermore, scMIR shows a strong generalization ability across various tasks without requiring task-specific fine-tuning. With its unique advantages, we envision scMIR may promote the standardization and automation of high-throughput phenotyping workflows through supporting various downstream analysis tasks.

[CV-192] StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design

链接: https://arxiv.org/abs/2607.22708
作者: Yin Wang,Haotian Hu,Jineng Han,Wentao Qiu,Zhenhua Ge,Liujian Tang,Fanyi Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the accuracy bar for OCR, screen understanding, visual question answering, and element grounding; on the other is the strict compute, memory, and power budget of mobile chips. Existing work either trades one for the other, or stops at simulation without real-device validation. We present StepX-Edge, a 0.9B-parameter on-device UI vision-language model that resolves this tension through three-layer co-design of architecture, training, and deployment. Architecturally, UI-aware Layered Visual Encoding (ULVE) and a Progressive Dimensionality Projection (PDP) connector target the extreme aspect ratios and fine-grained perception of screens, while standard full attention throughout ensures native compatibility with mainstream mobile NPU operators. For training, the five-stage StepX-Curriculum framework is designed around our observation of mutual-promotion effects among UI subtasks, so that all four capabilities grow synergistically under a tight parameter budget rather than interfering. For deployment, a module-wise differentiated two-stage PTQ-to-QAT quantization scheme keeps the post-quantization accuracy loss within 1%. StepX-Edge achieves the strongest overall UI understanding among =1B models, surpassing all 2B-2.3B baselines on ScreenQA (88.76 F1) and Chinese OCRBench v2 (57.25), and matching 1.3B-2.3B general VLMs on RefCOCO (92.0%) and OCRBench v1 (831) with far fewer parameters. After W4A16+KV8 quantization, the model runs stably on Snapdragon 8 Gen5 devices with ~0.84 s TTFT, 98 tok/s decode, and 1.4 GB peak memory. We will open-source the training data, the full training recipe, and the quantization deployment pipeline.

[CV-193] LowAux-RDNet: Low-Pass Residual Supervision with Scene-Balanced Real-World Training for Single-Image Reflection Removal

链接: https://arxiv.org/abs/2607.22707
作者: Jizhong Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 5 figures. Code: this https URL

点击查看摘要

Abstract:Single-image reflection removal aims to recover a clean transmission layer from one image captured through glass. We study an explicit decomposition pipeline built on RDNet and introduce LowAux, a training-only low-pass reflection auxiliary objective. The original residual target remains the main reflection supervision, while symmetrically filtered prediction and target provide a stable low-frequency constraint. We further incorporate scene-balanced real pairs from RRW to broaden real-scene coverage and improve cross-dataset generalization. To avoid evaluation discrepancies caused by model-specific resizing, padding, output quantization, and metric code, we build a unified public benchmark over CEILNet, Real20, Postcard, Objects, and Wild. Under the same evaluator, the proposed system obtains a five-dataset macro average of 27.546 dB PSNR, 0.9220 SSIM, 0.9751 NCC, and 0.004760 LMSE, achieving the highest macro-average PSNR, SSIM, and NCC and the lowest LMSE among the compared public checkpoints and internal variants. Per-dataset and qualitative analyses show that the main benefit is a more balanced performance across diverse reflection distributions, while clear semantic reflections in Postcard remain challenging.

[CV-194] EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations

链接: https://arxiv.org/abs/2607.22705
作者: Anuraag Gadehothur Karnam,Tarunesh Sathish
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations. Existing evaluations usually score segmentation, single-image factor prediction, or downstream accuracy, but these tests do not directly ask whether a per-object representation behaves correctly under a controlled semantic edit. We introduce EditCLEVR, a paired-scene intervention benchmark in which each example contains a before/after pair of CLEVR-style renders with the same object indices and scene layout, and either exactly one known attribute change on one known object or a no-edit re-render for drift measurement. The protocol includes probe-free diagnostics for representation-change localization and stability, together with probe-decoded semantic faithfulness metrics that test whether the predicted scene change matches the intended intervention across in-distribution and compositional out-of-distribution (OOD) suites, allowing code-space movement and decoded object-attribute correctness to be evaluated separately. We introduce the semantic metric Scene-Graph Intervention Accuracy (SGIA), which requires the full after-scene prediction to be correct and the only predicted before-to-after semantic change to be the intended object-factor edit. We also establish Delta-SGIA as a companion diagnostic that checks the single-site change pattern without requiring the full after-scene graph to be correct. Baseline evaluations on ground-truth-mask backbones, learned-slot models, SAM 2 + frozen-ViT models, and one mask-feature hybrid indicate that CoGenT-OOD-core degradation can persist under ground-truth instance masks, that mask source accounts for part but not all of native performance, and that locality or stability alone can overstate semantic faithfulness. Code is available at this https URL.

[CV-195] Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model

链接: https://arxiv.org/abs/2607.22704
作者: Xiao Wang,Hao Si,Qiang Chen,Yu-Xiang Zhang,Beihe Zhang,Jianhua Yang,Qingquan Yang,Dengdi Sun,Wanli Lyu,Guosheng Xu,Jin Tang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Nuclear fusion has made significant progress in recent years and is expected to become one of the most important pathways to addressing global energy challenges. This paper focuses on observing plasma using visible-light cameras, analyzing its spatio-temporal motion cues, and predicting the two-dimensional spatial distribution of light intensity, aiming to provide a foundational basis for future scientific experiments using deep neural networks. Specifically, we propose Delta-InvFormer, a novel backbone network centered on a differential Transformer. The key insight is that by taking consecutive video frames as input, we can better capture the dynamics of the plasma. Moreover, spatial and temporal differential self-attention effectively mitigates interference from noisy signals, ensuring high-quality feature extraction. These features are then fused into a compact and informative representation, which is fed into a decoder network to predict the distribution. Based on real experimental data collected from the Experimental Advanced Superconducting Tokamak (EAST) large-scale scientific facility, our results demonstrate that the proposed model not only significantly accelerates traditional methods for distribution prediction but also achieves competitive reconstruction accuracy. The source code of this paper will be released on this https URL

[CV-196] Histopathological Spectrum-Guided Prostate Stratification via Segmentation-Assisted Diagnostic Transformer

链接: https://arxiv.org/abs/2607.22703
作者: Leyang Li,Lihua Chen,Huangang Hu,Tianhang Hao,Hao Cheng,Xin Zhang,Qianru Sun,Bingxu Lu,Wenlong Yu,Feng Duan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 6 figures, and 4 tables. Code will be made publicly available in a future revision

点击查看摘要

Abstract:Prostate cancer diagnosis with multiparametric MRI (mpMRI) is commonly based on PI-RADS assessment or binary classification, which suffer from subjectivity and fail to capture clinically relevant pathological heterogeneity. To address this limitation, we construct a Prostate Cancer Histopathology Spectrum Dataset (PCa-HSD) and formulate a clinically meaningful four-class classification task, addressing the underrepresentation of benign lesions that are easily confounded with prostate cancer in existing datasets. We propose Language-guided Segmentation-assisted Diagnostic Transformer model (LSDT), which leverages zero-shot segmentation to provide anatomical priors and performs effective multi-modal slice fusion for classification. Our proposed method consistently improves accuracy across backbones, achieving the best average accuracy of 0.633 and JointRecall of 0.768 in five-fold cross-validation on a cohort of 344 patients. These results demonstrate that integrating pathology supervision and anatomical priors significantly enhances fine-grained prostate MRI classification and provides a more clinically relevant paradigm for risk stratification. Code will be made publicly available in a future revision.

[CV-197] MIME: Multimodal Interactive Motion Encoder WACV2027

链接: https://arxiv.org/abs/2607.22702
作者: Addison Zucek,Prerit Gupta,Kamila Kuatova,Aniket Bera
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Under review at WACV 2027

点击查看摘要

Abstract:Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. These settings require representations that align language with both individual actor dynamics and the relationships between actors. We introduce the Multimodal Interactive Motion Encoder (MIME), which, to our knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion. MIME captures individual and shared structure using stream based co-attention with explicit interaction features and curriculum based contrastive training. On Inter-X text-motion retrieval, MIME consistently outperforms early and late fusion baselines across gallery sizes, achieving a 12.8% relative improvement in text-to-motion R@1 at a 2,000-sample gallery. We further evaluate MIME as a frozen auxiliary prior within TIMotion and InterMask on the unseen InterHuman dataset. MIME improves semantic alignment metrics while maintaining comparable FID in TIMotion. These results show that interaction aware multimodal encoding improves multi person motion retrieval and transfers across datasets to support downstream motion generation.

[CV-198] FogDrive: A Multi-Modal Synthetic Driving Dataset for Perception under Graded Fog

链接: https://arxiv.org/abs/2607.22698
作者: Vansh Panwar
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Perception under adverse weather remains a critical bottleneck for reliable autonomous driving, yet existing benchmarks lack the systematic multi-modal alignments needed to evaluate robust sensor fusion. Real-world weather datasets suffer from uncontrolled collection and single-level, uncalibrated conditions, while synthetic alternatives either target camera-only restoration or lack the paired clean-and-foggy structure needed to benchmark “defog-then-detect” pipelines. We present FogDrive, a rigorously calibrated, multi-modal autonomous-driving dataset bridging data-centric engineering and robust machine learning. Built with the CARLA simulator, FogDrive contains 660 scenes (~133k fully annotated frames, 50:50 day/night) across four synchronized cameras (RGB, depth, semantic segmentation), a LiDAR and semantic-LiDAR pair, and front radar. Physically consistent fog is modeled independently on camera channels (Koschmieder model) and LiDAR channels (Beer-Lambert law) at three calibrated visibility densities (160m, 100m, 50m). Every scene ships in four matched variants (clean plus three graded fog levels) with cross-calibrated 2D and 3D bounding boxes. A semantic-segmentation-based quality audit over 8k images validates annotations at 95.1% precision and over 99% recall for vehicles within 40m. We establish baseline benchmarks with state-of-the-art architectures (TransFusion, BEVFusion, YOLOv8-m) across two paradigms: 3D multi-modal fusion and 2D image restoration. These yield critical data-centric insights: mixing multi-density fog during training tightens 3D bounding-box geometry without added data-scaling cost, while in 2D pipelines image-quality metrics (PSNR, SSIM) prove poor predictors of downstream detection performance. FogDrive will be fully open-sourced alongside our data-generation framework to accelerate robust, multi-modal research.

[CV-199] st-Time Coverag e: Test-Conditioned Data Curation for Deployment-Aware Learning

链接: https://arxiv.org/abs/2607.22697
作者: Nadine Chang,Maying Shen,Shizhe Diao,Jialiang Wang,Jingde Chen,Thomas Breuel,Pavlo Molchanov,Rafid Mahmood,Jose M. Alvarez
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution. However, standard data curation methods score training-side criteria rather than directly optimizing deployment match. We introduce TTCov (Test-Time Coverage), a data-level test-conditioned curation method that uses test-side information before training instead of updating model weights at inference. TTCov decomposes deployment-conditioned curation into coverage and distribution. To represent coverage, it builds a task Atlas, a collection of LLM-based atomic propositions (APs) describing deployment-relevant concepts, seeded from open task knowledge and expanded with unmatched APs extracted from unlabeled deployment samples. To represent distribution, it instantiates the matched deployment APs with their frequencies, yielding a Knowledge Atlas (K-Atlas) that operationalizes the deployment distribution as a curation target. TTCov then selects a budgeted training set whose deployment APs distribution approximates this target. We apply TTCov towards autonomous driving (AD), keeping adaptation off the inference path while selecting data with greater deployment-relevant coverage, closer K-Atlas matching, and stronger downstream end-to-end driving performance than data-curation baselines, including seamless adaptability to novel domains via city-to-city expansion.

[CV-200] MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion

链接: https://arxiv.org/abs/2607.22696
作者: Jiacheng Liu,Jason Liu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget of a single workstation. A 100 billion-plus parameter DiT easily requires over a terabyte of persistent state, while naive spatiotemporal self-attention grows quadratically in sequence length. These two walls – parameter memory and activation memory – prevent researchers from adapting massive generative models without large GPU clusters. We revisit this problem from a systems perspective and introduce MegaSlide-DiT, a prototype that demonstrates how a pre-trained 105B DiT can be adapted on a single H200 GPU with 1.5 TB of host RAM. Our key insight is that the GPU need not own the model state: all persistent weights, master weights and optimizer moments remain in host memory, while only transient shards are streamed to the GPU on demand. Simultaneously, we replace quadratic global attention with 3D Deformable Slide Attention (3D-DSA), a motion-adaptive local attention operator that reduces both memory and computational complexity to linear in the sequence length. We report detailed memory accounting, execution traces and evaluation results to substantiate our design. MegaSlide-DiT does not claim to train a 105B model from scratch on a single GPU, nor does it magically solve bandwidth limits; rather, it offers a pragmatic path for full-parameter adaptation of massive video diffusion models on high-end workstations.

[CV-201] DINOv3-MIL: Per-Kidney Multi-Label Tumour and Cyst Detection from Foundation-Model Patch Tokens on KiTS23

链接: https://arxiv.org/abs/2607.22687
作者: Vishalakshi M,Sahil Sharma,Pramod Kumar P
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at MIUA 2026 (poster). To appear in Frontiers in Medical Technology

点击查看摘要

Abstract:Foundation vision models trained on natural images transfer to medical tasks without domain pre-training, but volumetric classification requires aggregating tens of thousands of patch tokens per study, and the aggregator constrains how the resulting model can be interpreted. We compare three aggregators on identical frozen DINOv3 ViT-H/16+ features for renal tumour/cyst detection on KiTS23 (966 kidneys; n=97 test): a CLS-token linear probe, gated attention multiple instance learning (MIL) over 55,296 patch tokens, and a prototype head following ProtoViT. Attention MIL achieves the highest AUROC for tumour (0.74, 95% CI 0.64-0.83) and cyst (0.80, 0.70-0.88), with attention enriched 7.5-9.8x over chance within annotated lesions. The prototype head does not transfer to cyst detection (AUROC 0.51), exposing an interpretability-performance trade-off at this token scale.

[CV-202] URHead: A Unified UV-Space Representation for Joint Mesh-3DGS Optimization in Head Avatars ECCV2026

链接: https://arxiv.org/abs/2607.22673
作者: Seonghak Lee,Junhee Cho,Jisoo Park,Min-Gyu Park,Jongmin Lee,Ju Hong Yoon,Junseok Kwon
类目: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
备注: Project page/code: this https URL , Accepted to ECCV 2026

点击查看摘要

Abstract:We present URHead, a unified representation for high-fidelity and animatable head avatars that fundamentally redefines mesh-Gaussian integration. While mesh-based methods offer precise geometric control but lack photorealistic detail, and Gaussian-based approaches achieve photorealism but suffer from poor structural consistency, existing hybrid solutions fail to fully leverage their complementary strengths. Our key contribution is a UV-space unification where both representations share a common UV parameterization. Through joint optimization with adaptive gaussian sampling, our method automatically learns to disentangle and allocate appropriate roles to each component. URHead maintains full parametric controllability while preserving subject-specific details, and outperforms existing state-of-the-art methods in reconstruction quality and animation consistency.

[CV-203] DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

链接: https://arxiv.org/abs/2607.22644
作者: Mohammed Yousif,Prabhjot Singh,Arjun Pankajakshan,Madhu Reddiboina
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 9 pages, 4 tables, 2 figures

点击查看摘要

Abstract:Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or type. This leads to inefficient use of compute and human resources: simple documents are over-processed while difficult ones may not receive enough scrutiny. We introduce DocHRL, a hierarchical reinforcement learning framework that learns to adaptively and dynamically select the most cost-effective classification policy on a per-document basis. DocHRL formulates document classification as a sequential decision problem with a two-level policy hierarchy: a top-level policy selects among broad options (vision classifiers, LLMs, OCR, and human-in-the-loop review), while option-specific sub-policies choose the concrete model or tool to invoke. The reward signal is the negative total expected cost, which captures inference cost, cost of misclassification, and cost of human labelling. Trained with Proximal Policy Optimisation on the RVL-CDIP benchmark, DocHRL achieves a macro F1 of 0.973 across 16 document classes while reducing average per-document cost to 2.74 normalised units compared to substantially higher costs incurred by fixed standalone classifiers. Our results demonstrate that cost-aware reinforcement learning can simultaneously improve classification performance and operational efficiency in document understanding systems.

[CV-204] Reason Before You Retrieve: Agent ic Planning for Multi-modal RAG

链接: https://arxiv.org/abs/2607.22643
作者: Tianyu Yang,Shir Simon,Zhenzhen Li,Minhao Cheng,Xiangliang Zhang
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space. This design often struggles with two key challenges: the retrieval target is under-specified because the question intent must be grounded to the correct visual referent, and the search space is weakly structured, forcing semantically distinct evidence to compete in a single global ranking step. We propose MM-R2, a multimodal agentic retrieval framework that reasons before retrieval by explicitly modeling both what to retrieve and where to search. MM-R2 first constructs an intent-grounded retrieval state from the image-question pair, capturing the information need, grounded referent, and retrieval constraints. It then performs retrieval over a structured KnowledgeMap, where the agent selects relevant retrieval units before issuing grounded queries within them. To enable this capability, we build MM-R2-Traj, a large-scale trajectory dataset of multi-step retrieval processes, and adopt a two-stage post-training strategy with supervised fine-tuning and GRPO. Experiments on Infoseek and Encyclopedic VQA datasets show that MM-R2 substantially outperforms strong baselines on answer accuracy while also yielding more interpretable and verifiable retrieval trajectories.

[CV-205] Concept-based Visual Counterfactual Explanations with Diffusion Models

链接: https://arxiv.org/abs/2607.22544
作者: Yassine Oueslati,Daniil Kirilenko,Martin Gjoreski,Marc Langheinrich
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the 4th World Conference on eXplainable Artificial Intelligence (XAI 2026)

点击查看摘要

Abstract:Visual counterfactual explanations aim to answer “what minimal change to this image would flip the model’s prediction?”, and are increasingly important as vision models are deployed in safety-critical domains (e.g., medicine). Existing diffusion-based methods can produce realistic edits, but they rely on external classifiers that must work reliably on noisy images, which makes them fragile and hard to deploy for robust explanations. We introduce C-VCE, a new diffusion framework that builds the classifier directly into the generative model via a concept bottleneck layer, so that counterfactuals are guided by human-interpretable features (concepts) instead of a separate noise robust classifier that works with pixel-level edits. Our model lets users to toggle on/off semantic concepts during sampling, then minimally adjusts relevant image regions, while preserving the rest of the image, respecting feature correlations. To keep edits small and controlled, we add a simple probabilistic regularizer that balances “change the prediction” against “stay close to the original”, plus a gradient-based mask that confines modifications to the most relevant regions. On benchmarks such as CelebA, C-VCE matches or improves flip rates while producing counterfactuals that are visually closer to the input and less distorted than baselines that depend on separate noisy-image classifiers. These properties make C-VCE a practical tool for vision systems where users need concrete “what-if” images without having to trust an additional, noise-robust classifier. More broadly, our results suggest that exposing and controlling an internal concept layer is a promising way to make powerful generative models easier to understand and safer to use.

[CV-206] Face Recognition with Machine Learning in OpenCV_ Fusion of the results with the Localization Data of an Acoustic Camera for Speaker Identification

链接: https://arxiv.org/abs/1707.00835
作者: Johannes Reschke,Armin Sehr
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Applied Research Conference 2017 (Munich)

点击查看摘要

Abstract:This contribution gives an overview of face recogni-tion algorithms, their implementation and practical uses. First, a training set of different persons’ faces has to be collected and used to train a face recognizer. The resulting face model can be utilized to classify people in specific individuals or unknowns. After tracking the recognized face and estimating the acoustic sound source’s position, both can be combined to give detailed information about possible speakers and if they are talking or not. This leads to a precise real-time description of the situation, which can be used for further applications, e.g. for multi-channel speech enhancement by adaptive beamformers.

[CV-207] A Modern ConvNet for Solar Filament Detection

链接: https://arxiv.org/abs/2607.24525
作者: J. R. Hu,Q. Hao,Z. Zheng,P. F. Chen,C. Li,Y. Meng
类目: olar and Stellar Astrophysics (astro-ph.SR); Computer Vision and Pattern Recognition (cs.CV)
备注: 21 pages, 4 figures, accepted for publication in RAA

点击查看摘要

Abstract:Automated solar filament detection using deep learning faces several challenges. Semantic segmentation of solar filaments is a complicated multiscale feature extraction task with long-tail distribution. Furthermore, a large-scale, highly complete, and finely detailed dataset has become mandatory for providing abundant information. To address these challenges, we present a series of machine learning approaches to develop a solar filament detection workflow that performs superbly. First, we manually annotated a small-scale solar filament dataset based on H \alpha spectra called MHAS. Next, we developed the Multiscale ORiented DENdritic (MORDEN) model, a semantic segmentation model focusing on multiscale feature extraction. We also introduced the Dense Conditional Random Field (DenseCRF) and Density-Based Spatial Clustering of Applications with Noise (DBSCAN) methods for post-processing. Using the proposed workflow, we generated a large-scale, high-quality dataset called AHAS. Experimental results demonstrate that MORDEN outperforms several existing solar filament semantic segmentation models with open access. DenseCRF has been demonstrated to effectively capture fine edge details. We also evaluated the effects of data scaling and the reliability of DBSCAN and found that both approaches yield satisfactory performance. Multiple visualization results substantiate our quantitative findings. Our work provides a foundation for maximizing the potential of deep learning models for solar filament detection.

[CV-208] A Controlled Visual-Backbone Benchmark for Multimodal Short-Term Solar Irradiance Forecasting

链接: https://arxiv.org/abs/2607.23633
作者: Oshadha Samarakoon,Dushan Herath,Ishara Ranmandala,Dilshara Herath,Roshan Godaliyadda,Parakrama Ekanayake,Vijitha Herath
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages, 3 figures, Moratuwa Engineering Research Conference 2026 (MERCon 2026)

点击查看摘要

Abstract:Sky-image irradiance studies often compare forecasting systems in which the image encoder, temporal model, fusion block, target definition, and training recipe all change together. We use a narrower protocol: the multimodal forecasting pipeline is fixed, and only the visual backbone is varied. The shared setup keeps preprocessing, clear-sky-index normalization, weather-history encoding, fusion, regression head, loss, optimizer schedule, seed, and chronological split policy unchanged. We compare ConvNeXt, Swin Transformer, VMamba, Spatial Mamba, and MambaVision backbones for 10min-ahead forecasting on Folsom and a strict matched NREL split. Forecast skill is measured against clear-sky-index smart persistence, and temporal-only rows are reported as weather-history diagnostics rather than as the main ranking criterion. On the Folsom strict split, all evaluated visual-backbone runs improve over smart persistence. In the evaluated single-seed strict runs, VMamba Small and Swin Base reach matched Folsom RMSE values of 65.39 W/m^2 and 65.50 W/m^2; the temporal-only diagnostic reaches 69.51 W/m^2. On the 313-sample NREL strict split, smart persistence remains strongest at 17.48 W/m^2, while the lowest visual RMSE is obtained by Swin Tiny at 23.76 W/m^2. These results provide a reproducible encoder comparison under one fixed multimodal operating point rather than establishing architecture-level dominance, statistically resolved ranking, or fully optimized forecasting performance. Code available here: this https URL

[CV-209] Segmentation Robustness and Predictive Utility in Glioblastoma Radiomics: Evidence for a Trade-off in Survival Modelling

链接: https://arxiv.org/abs/2607.23626
作者: Mariya Miteva,Maria Nisheva-Pavlova
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted as a full paper at HCist 2026 and for publication in Procedia Computer Science

点击查看摘要

Abstract:Radiomic biomarkers derived from magnetic resonance imaging (MRI) have been widely investigated as non-invasive tools for tumor characterization and prognostic modeling in glioblastoma (GBM). However, their clinical translation remains limited, in part due to sensitivity to tumor segmentation variability. In this study, we systematically investigate the relationship between feature robustness and predictive utility in GBM survival modeling using the University of Pennsylvania Glioblastoma Imaging, Genomics, and Radiomics (UPENN-GBM) cohort. A total of 4,752 radiomic features were obtained from multiparametric MRI across three tumor subregions: enhancing tumor (ET), peritumoral edema (ED), and necrotic core (NC). Feature robustness was quantified using the intraclass correlation coefficient (ICC) based on the automatic and expert-refined segmentation versions. Among features with valid ICC estimates, 48.1% were classified as robust. Survival prediction was evaluated using cross-validation with Coxnet, Random Survival Forest, and Gradient Boosting Survival Analysis models. In this cohort, radiomic feature inclusion showed no consistent improvement over the clinical baseline, and robustness filtering produced no detectable performance gain. Model-selected features were less robust than the overall feature pool, indicating a lack of enrichment for robustness. These findings suggest that robustness alone is not a reliable criterion for feature selection in radiomics-based survival modelling.

[CV-210] Direction-adaptive Mamba: Spatial-Frequency Dual-Domain Collaborative Learning for PolSAR Image Classification

链接: https://arxiv.org/abs/2607.23464
作者: Junfei Shi,Yu Cheng,Haojia Zhang,Wenqiang Hua,Junhuai Li,Maoguo Gong
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Deep learning dominates polarimetric synthetic aperture radar (PolSAR) image classification, with Mamba architectures serving as favorable backbones due to linear complexity and strong global modeling capacity. However, existing PolSAR Mamba methods have two critical flaws: pure spatial processing discards fine-grained edges and textures, and fixed scanning patterns fail to model direction-variant anisotropic scattering and weak boundaries essential for PolSAR physical analysis. This work proposes DA-Mamba, a direction-adaptive Mamba framework with dual-domain collaborative learning for PolSAR classification. Equipped with an edge-aligned direction-adaptive scanning scheme, DA-Mamba captures long-range spatial dependencies and accurate boundary details. It adopts the Non-Subsampled Contourlet Transform (NSCT) to separate PolSAR data into low-frequency global components and multi-directional high-frequency subbands, extracting anisotropic structural features from high-frequency information while preserving global context via low-frequency branches. A dual-domain collaborative learning module further integrates spatial scattering and frequency-domain representations to strengthen feature discriminability. Evaluated on three real-world PolSAR datasets, DA-Mamba surpasses state-of-the-art methods, verifying the efficacy of the proposed adaptive scanning and dual-domain fusion designs. Code will be publicly available.

[CV-211] Investigating the Visual Cues of CNNs for Vascular Segmentation: A Case Study in Microscopy and Fundus Imaging

链接: https://arxiv.org/abs/2607.23371
作者: Weslley dos Santos Silva,Cesar Henrique Comin
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Vascular segmentation is a standard procedure for clinical diagnosis, yet the specific visual features determining model decisions remain poorly understood. This paper investigates the visual cues Convolutional Neural Networks (CNNs) use to segment blood vessels across two distinct imaging domains: fluorescence microscopy and retinal fundus photography. We employ a series of experiments to quantify the influence of shape, texture, and receptive field on segmentation performance. First, we isolate texture and intensity by evaluating performance on patches subjected to pixel shuffling and normalization. Second, we assess global shape relevance by training models on sparse contours and centerlines. Lastly, we quantify the required spatial context by systematically varying the network’s theoretical and effective receptive fields. Within the scope of the evaluated datasets, we found that pixel intensity is more relevant than texture, though networks maintain surprisingly high accuracy even when both cues are removed. Furthermore, CNNs struggle to extrapolate full vessel geometry from shape cues alone, typically relying on a relatively small effective receptive field of around 20 pixels, though global context provides a modest benefit for fundus images. While specific to the modalities studied, this methodology offers a quantitative foundation to audit and refine deep learning systems in vascular imaging.

[CV-212] Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression

链接: https://arxiv.org/abs/2607.23366
作者: Manikanta Kotthapalli,Banafsheh Rekabdar
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:

点击查看摘要

Abstract:Learned video codecs based on continuous latent representations typically require resolution-specific retraining or rate-distortion (RD) recalibration when scaling to new spatial resolutions, because entropy models and Lagrangian weights are tightly coupled to the operating point. We investigate whether hierarchical discrete latent codecs exhibit the same sensitivity. Using a controlled empirical study of MS-VQ-VAE video compression across codebook sizes K \in \128,256,512,1024\ and resolutions 64\times64 , 128\times128 , and 256\times256 on UCF101, we show that perceptual quality (LPIPS) depends strongly on codebook capacity but only negligibly on spatial resolution. Fitting a log-linear model Q(K,r) = \alpha\log_2 K + \beta\log_2 r + \gamma to all 12 operating points yields \alpha=-0.0094 ( t=-6.6 , p0.001 ) and \beta=-0.0009 ( t=-0.43 , p=0.68 , not significant), with R^2=0.82 . Codebook capacity is therefore roughly 10\times more influential than spatial resolution per log-unit increase. In parallel, bottom-level entropy efficiency \eta=H(z)/\log_2 K remains stable or improves with resolution (84-87% at 64\times64 ; 92-94% at 256\times256 ), confirming that larger spatial grids are utilized more efficiently rather than less. Across all resolutions and codebook sizes, our models outperform H.264 on LPIPS at matched or lower bitrate, with gains of 25-52% at 128\times128 and 21-37% over H.265 at 256\times256 . These findings suggest that codebook size K , not spatial resolution, is the dominant design variable governing perceptual compression quality in hierarchical discrete video codecs – a property that may simplify multi-resolution deployment and inform the design of scalable discrete tokenizers for generative video models.

[CV-213] rainable Nonexpansive Denoisers for Contractive Image Reconstruction ICML2026

链接: https://arxiv.org/abs/2607.23347
作者: Arghya Sinha,Aditya Banerjee,Trishit Mukherjee,Kunal N. Chaudhury
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: accepted at ICML 2026

点击查看摘要

Abstract:Trainable denoisers with Lipschitz control have become central to convergent image reconstruction. However, training neural networks that simultaneously offer strong denoising performance and global Lipschitz guarantees is challenging. Existing approaches enforce Lipschitz control only empirically, providing no guarantees beyond the training data. In this work, we show that by exploiting the action of permutations on the image lattice, we can constrain a neural architecture that is globally nonexpansive (Lipschitz bound \leqslant 1 ). We integrate the proposed denoiser with forward imaging operators to develop a reconstruction mechanism that is provably contractive and therefore globally convergent. Experiments on standard inverse problems, such as superresolution and deblurring, demonstrate that our reconstruction performance is competitive with softly constrained baselines while providing Lipschitz guarantees.

[CV-214] Patient-Agnostic Synthetic Pretraining for Efficient Patient-Specific Intraoperative 2D/3D Registration

链接: https://arxiv.org/abs/2607.23343
作者: Minheng Chen,Youyong Kong
类目: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Intraoperative 2D/3D registration aligns preoperative CT volumes with intraoperative X-ray or fluoroscopic images and is essential for image-guided interventions. Recent learning-based and differentiable registration methods have shown promising accuracy, especially in patient-specific settings where abundant digitally reconstructed radiographs (DRRs) can be synthesized from the target CT. However, training a separate patient-specific model from scratch for every new patient is computationally inefficient and limits practical deployment. In this work, we propose an efficient patient-specific 2D/3D registration framework based on patient-agnostic synthetic pretraining and spherical similarity learning. The model is first pretrained on synthetic DRRs generated from multiple CT volumes to learn transferable pose-sensitive representations, and is then adapted to a new patient using only a limited number of synthetic projections from the target CT. To improve synthetic-to-real robustness without requiring anatomical labels, we introduce a segmentation-free domain randomization strategy that perturbs image intensity, projection physics, field-of-view, occlusion, and fluoroscopic artifacts. The adapted model provides an initial pose estimate, which is further refined using spherical similarity learning and differentiable Levenberg-Marquardt optimization. Experiments on multiple anatomical datasets evaluate whether patient-agnostic synthetic pretraining can improve the efficiency of patient-specific registration, with particular focus on the trade-off between adaptation cost and registration accuracy. The results demonstrate that patient-agnostic synthetic pretraining can significantly reduce patient-specific training requirements while preserving accurate intraoperative 2D/3D registration.

[CV-215] Stabilizing Deep Reconstruction Operators with Contractive Anchoring ECCV2026

链接: https://arxiv.org/abs/2607.23341
作者: Arghya Sinha,Trishit Mukherjee,Kunal N. Chaudhury
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: accepted at ECCV 2026

点击查看摘要

Abstract:Pretrained deep denoisers can be used to solve a wide range of model-based image reconstruction tasks via Plug-and-Play (PnP) and Regularization-by-Denoising (RED) algorithms, without retraining per task. These denoisers are trained only for single-step denoising. Using them as Image Reconstruction (IR) regularizers in an iterative process can destabilize reconstruction. A common failure mode is the peak-and-collapse behaviour: metrics such as PSNR improve for early iterations and then abruptly degrade, making these algorithms unreliable in practice. We propose a data-driven stabilization framework that (i) formalizes this instability of any IR operator through a local quantity and (ii) prevents collapse by regularizing this quantity adaptively, requiring no retraining or modification of the given pretrained network. Our key idea is to control the potentially unstable IR operator with a contractive operator whose stable iterates act as an anchor and prevent collapse. We further introduce an efficient family of trainable contractive operators that serve as strong anchors while remaining lightweight. Extensive experiments across proximal algorithms, denoiser architectures, noise levels, and imaging tasks show consistent, collapse-free performance and improved reliability of PnP and RED reconstruction.

[CV-216] A Reference-Free Framework for Evaluating Single-Frame ISP Pipelines

链接: https://arxiv.org/abs/2607.23321
作者: Yujin Cho,Sira Ferradans,Jean-Michel Morel,Gabriele Facciolo,Thomas Eboli
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 8 figures, 11 tables

点击查看摘要

Abstract:Evaluating camera image signal processing (ISP) pipelines requires measuring low-level artifacts introduced by operations such as denoising, demosaicing, tone mapping, and compression. Blind image quality assessment (IQA) techniques can grade visual quality without a reference, but they typically focus on semantic and high-level visual cues or human perceptual scores rather than the low-level image-processing artifacts introduced by camera pipelines. In contrast, full-reference metrics such as PSNR and SSIM measure pixel-level differences and structural similarity, while LPIPS measures perceptual similarity in deep feature space. However, these metrics require perfectly aligned image pairs, which are difficult to collect in practical settings. We propose a reference-free learning framework that estimates full-reference image quality metrics from a processed sRGB image and its ISO metadata. Our method predicts a proxy sRGB reference, which is then compared with the processed image to compute PSNR, SSIM, and LPIPS in their standard full-reference form. Our experiments show that the proxy-reference model can be learned from synthetic data and applied to real camera data. We further show that lightweight LoRA fine-tuning enables efficient adaptation when ISP components or pipeline configurations are changed. The proposed method outperforms direct metric regression in estimating metric values and achieves higher agreement with full-reference rankings than conventional blind IQA methods. These results demonstrate the feasibility of reference-free estimation of full-reference metrics for practical camera-pipeline evaluation.

[CV-217] PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation

链接: https://arxiv.org/abs/2607.22963
作者: Fan Zhang,Xuanting Wu,Fei Ma,Qiang Yin,Yuxin Hu
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages,15 images

点击查看摘要

Abstract:Synthetic aperture radar (SAR) image generation can mitigate data scarcity, but controllablegeneration under sparse observation angles remains difficult. Recent SAR generative studies im-prove texture realism, yet explicit geometry-aware control is still limited. This paper studiesthe focused and verifiable setting of intermediate-azimuth completion: 3D-model-derived geo-metric priors guide a diffusion model to synthesize the views missing from sparse-angle trainingdata. GeoDiff-SAR constructs a lightweight multi-bounce ray-tracing prior, encodes the result-ing point cloud, and fuses it with text conditioning while adapting Stable Diffusion 3.5 Mediumthrough low-rank adaptation. On a real four-category aircraft dataset, GeoDiff-SAR reaches anSSIM of 0.812 and azimuth consistency of 0.940, compared with 0.738 and 0.782 for the text-conditioned SD3.5 Medium baseline. The same sparse-angle protocol on five MSTAR vehicleclasses yields an SSIM of 0.878 and azimuth consistency of 0.917. These results support theconclusion that a lightweight 3D geometric prior improves viewpoint adherence for controllableSAR generation; it is intended as generation guidance rather than high-fidelity electromagneticreconstruction.

[CV-218] Agent ic Autoresearch for CT Reconstruction

链接: https://arxiv.org/abs/2607.22824
作者: Andreas Maier,Lucas Kachelriess,Siming Bayer,Yixing Huang,Yan Xia,Amber Simpson,Moritz Zaiss
类目: Medical Physics (physics.med-ph); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 16 pages

点击查看摘要

Abstract:Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a large language model (LLM) agent can do the labor of reconstruction research on its own, and whether a ranking measured on ideal data predicts behavior under realistic noise. We built an agentic loop: the agent edits a solver, runs a short cluster job, reads one frozen metric, and revises. The metric is a calibrated headroom score against the FBP baseline, inside the field of view; every method shares the same differentiable fan-beam projector. We benchmarked 26 methods on Mayo low-dose CT (noise-limited) and a 128-view sparse-view breast task from the noiseless DL-Sparse-View Challenge, with validation-selected iterations scored on a held-out test set. Every trained breast model was then re-scored on noisy inputs (I_0 = 10^5 photons) without retraining, and separately retrained on matched noise. The agent independently implemented, tuned, and benchmarked all 26 methods, and recombined them into a compact solver of 969 parameters that ties the top Mayo tier at the 1% level using 0.4% of the champion’s parameters. Benchmarking gives a tier of statistically indistinguishable top methods, not one winner. Mild input noise nearly inverts the breast ranking: the noiseless champion (a supervised image denoiser, hr 0.89) collapses to 0.00, while a learned primal-dual method rises to champion (0.72 to 0.93). An ideal-data leaderboard therefore does not predict robustness. The inversion is a transfer effect, not a permanent deficit: retraining on matched noise restores much of the clean ranking (Spearman rho 0.04 to 0.61). Noise is only the easiest confounder in an open-ended set (beam hardening, scatter, anatomy, disease), so no single-factor challenge certifies generality. Benchmarks should model a broad spectrum of realistic factors at once. Comments: 16 pages Subjects: Medical Physics (physics.med-ph); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.22824 [physics.med-ph] (or arXiv:2607.22824v1 [physics.med-ph] for this version) https://doi.org/10.48550/arXiv.2607.22824 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Andreas Maier [view email] [v1] Fri, 24 Jul 2026 18:03:21 UTC (1,830 KB)

[CV-219] Learning Dense 2D-3D Correspondence for X-ray-to-CT Registration of Knee Bones

链接: https://arxiv.org/abs/2607.22803
作者: Rembert Daems,Jonas Grammens,Caro Roten,Andrew Meyer,Thomas Luyckx,Matthias Verstraete
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recovering the 6-DoF pose of the knee bones from a plain radiograph, given the patient’s segmented pre-operative CT, turns a routine low-dose image into a quantitative measurement of joint geometry, without the added dose of a repeat CT or a fixed biplanar rig. Classic solutions align a rendered bone silhouette to image edges; recent alternatives refine pose by backpropagating an image-similarity loss through a differentiable X-ray renderer. Both operate one patient at a time and are fragile under a single view. Silhouettes are depth-ambiguous, and differentiable-rendering refinement has a narrow capture range at substantial per-iteration cost. We instead learn an amortized, subject-agnostic dense 2D-3D correspondence, supervised solely by projection geometry. One shared-weight model per bone, trained across 758 patients, registers patients unseen during training. The pose then follows in closed form from a global, initialization-free, render-free PnP+RANSAC solve. Because X-ray formation is transmissive, our correspondence target is transmission-aware rather than tied to a single surface. Though trained only to register, the representation is anatomically semantic: a simple classifier reads a landmark’s anatomical region from its embedding across held-out patients, and the same features separate the knee’s bones into a 2D-3D-consistent identity learned without any bone label. On a large single-institution cohort the model generalizes well to held-out patients.

[CV-220] DY-LUT: Depth-Aware YCbCr Lookup Tables for Real-Time Underwater Image Enhancement

链接: https://arxiv.org/abs/2607.22801
作者: Cunhao Zhu,Xiangtao Kong,Dongliang Xu,Zhiheng Zhang,Tianyu Wang,Yue Yao
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Underwater image enhancement is challenged by spatially non-uniform, wavelength-dependent attenuation. Propagation distance and wavelength govern this degradation, while YCbCr separates luminance from chrominance for restoration. We propose DY-LUT, a depth-aware YCbCr lookup-table framework for real-time enhancement. A dual-branch encoder predicts image-level fusion weights and a joint pair of pixel-wise degradation indices from image and depth features. These quantities condition learnable 4D LUTs, followed by lightweight local refinement. DY-LUT preserves traditional LUT efficiency while enabling depth-conditioned, spatially adaptive restoration. With externally supplied depth, its 3.56M-parameter enhancement network achieves competitive quality on UIEB-90 and LSUI and runs 9 – 304\times faster than representative high-capacity baselines. Adaptive inference further maintains real-time performance ( \sim7 ms) for 4K UIQAD images. DY-LUT also benefits downstream detection and feature matching. Ablations show that YCbCr is a more effective basis than RGB for depth-conditioned lookup, while the jointly learned indices further improve adaptive querying. These results provide a physically grounded route to efficient UIE on practical platforms.

[CV-221] Small Bias-Free Blind and Convolutional Denoiser: A compact ConvNeXt U-Net for blind Gaussian color-image denoising

链接: https://arxiv.org/abs/2607.22793
作者: Nikolas Markou
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 1 figure

点击查看摘要

Abstract:We describe and evaluate BF-ConvUNeXt, a compact bias-free ConvNeXt U-Net for blind additive-white-Gaussian-noise color image denoising, combining four existing ingredients so a single property survives end to end: a frozen depthwise Gabor stem (oriented band-pass, zero trainable parameters), a Laplacian-pyramid encoder routing the high-frequency residual into each skip connection, a ConvNeXt-V1 U-Net body, and bias-free construction throughout (no additive bias, linear head, LeakyReLU, variance-only batch norm). Together these make the 0.82M-parameter network exactly degree-1 homogeneous at inference, D(alpha y) = alpha D(y), licensing a Miyasawa/Tweedie score reading of the residual and blind generalization across noise levels from one model. We train a single blind model on a noise-sigma curriculum (sigma approximately 6.4 to 64, 0-255 scale); it extrapolates past that ceiling without a cliff, degrading smoothly to 22.8 dB at sigma=150 and 20.0 dB at sigma=200. Evaluated unchanged on four standard color sets (CBSD68, Kodak24, McMaster, Urban100) at sigma in 15,25,50, it matches or beats DnCNN and FFDNet on every set and level, averaging about +0.7 dB over DnCNN. Against heavyweight CNN/transformer state of the art it trails by a small margin (roughly 0.3-1.7 dB depending on set) at 1/15 to 1/39 of their parameters. The homogeneity is inference-only and checkpoint-specific, and the learned residual is a local, not global, score (non-conservative Jacobian), so plug-and-play/RED guarantees do not transfer; it still drives stochastic sampling and linear inverse problems (inpainting, super-resolution, deblurring, compressive sensing).

[CV-222] Learning-based Hierarchical Tracheal Anatomy Understanding from Sparse Surgical Demonstration Annotations for Ultrasound Robots

链接: https://arxiv.org/abs/2607.22789
作者: Hiu Ching Cheung,Wenchao Yue,Zhengran Han,Mingcong Chen,Guanglin Cao,Hongbin Liu,Hongliang Ren
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 17 pages, 8 figures, accepted at the 2026 International Conference on Cyborg and Bionic Systems (2026ICCBS)

点击查看摘要

Abstract:Tracheostomy requires precise localization of the tracheal incision site; however, conventional manual palpation is subjective and often unreliable, while ultrasound utility remains operator-dependent. This work presents a learning-based framework for hierarchical tracheal anatomy understanding, designed specifically for ultrasound-guided robotic systems. We propose a two-stage perception pipeline integrating a YOLOv8n localization backbone with a sparse, prompt-optimized SAM2 decoder to achieve high-fidelity segmentation from sparse surgical annotations. Our hybrid training strategy, bridging curated laboratory data with unconstrained sequences, ensures clinical robustness. Experimental benchmarks demonstrate that this decoupled architecture effectively balances generalization, precision, and efficiency. The YOLOv8n and SAM2 framework achieves a consistent Mean Dice Similarity Coefficient (DSC) of 0.777 across both controlled and generalized domains. This significantly outperforms U-Net baselines, which often suffer from anatomical fragmentation and performance degradation (Generalization DSC \le 0.494). By constraining mask decoding to targeted, sparse regions of interest, our model achieves a throughput of 6.92 FPS, which is vital for closed-loop robotic teleoperation. This study confirms that a robust hierarchical understanding of tracheal anatomy can be derived by coupling lightweight localization with foundation-scale visual models. Our framework establishes a scalable foundation for standardized, autonomous surgical assistance, effectively navigating the variability of real-world ultrasound to enhance the safety and precision of robotic-assisted tracheostomy.

[CV-223] JPEG AIC2026: A large-scale dataset for fine-grained assessment of image coding

链接: https://arxiv.org/abs/2607.22783
作者: Mohsen Jenadeleh,Jon Sneyers,João Ascenso,Thomas Richter,Alexander Karabutov,Panqi Jia,Elena Alshina,Osamu Watanabe,António Pinheiro,Touradj Ebrahimi,Dietmar Saupe
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Recent advances in conventional and learning-based image coding have increased the demand for benchmark datasets that support fine-grained assessment of compressed image quality, particularly for learning-based image compression methods. This paper introduces Assessment of Image Coding 2026 (AIC2026), a large-scale dataset for high-fidelity image compression containing 70 source images selected from 2,787 candidates using semantic clustering, inter-metric disagreement among objective image quality assessment (IQA) methods, and manual inspection and refinement. The dataset covers a wide range of compression artifacts produced by eight conventional and four learning-based codecs across 17 coding configurations. Each source image is encoded using seven codecs. For each source-codec pair, decoded images are provided at 20 perceptually spaced distortion levels, corresponding approximately to 0.2-4.0 just-noticeable difference (JND) units using the ColorVideoVDP (CVVDP) metric for distortion estimation, yielding 9,618 distorted images. This fine-grained sampling enables analysis of rate-distortion behavior and objective metric evaluation for subtle quality differences across a wide range of compression artifacts. We report an extensive objective analysis using 24 conventional and 12 learning-based IQA methods. The results show substantial disagreement among current IQA methods for fine-grained quality differences, particularly for artifacts introduced by learning-based codecs. The complete dataset is publicly available at this https URL.

[CV-224] Metric Surface Reconstruction of Neurosurgical Scenes from Monocular Operating Microscope Images and Microscope Pose

链接: https://arxiv.org/abs/2607.22773
作者: Thomas Bucher,Didier Neuenschwander,Thomas Petutschnigg,Michael Murek,David Bervini,Andreas Raabe,Manuela Eugster
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Medical Physics (physics.med-ph)
备注:

点击查看摘要

Abstract:Objective: We evaluated whether metric 3D geometry of neurosurgical operative exposure can be recovered from standard monocular operating-microscope images combined with microscope pose data. Methods: In a phantom-based laboratory study, two aneurysm training phantoms were imaged with a ZEISS Pentero 800 microscope integrated with Brainlab Cranial Navigation. Microscope images from the standard composite video output were stored with synchronous microscope poses. After intrinsic and extrinsic calibration, depth was estimated with the pretrained Depth Anything 3 model without task-specific fine-tuning. Fused point clouds were converted to meshes using Poisson surface reconstruction. Reconstructions were compared with reference surfaces from structured-light scanning and fine-slice CT. Results: For phantom A, representing a deeper surgical corridor, reconstruction accuracy ranged from 1.95 \pm 1.70 mm to 2.33 \pm 2.15 mm. For phantom B, representing a directly exposed surface, accuracy ranged from 1.02 \pm 0.93 mm to 1.52 \pm 1.21 mm. Larger image sets mainly improved completeness, while accuracy remained within a narrower range. Corridor analysis showed preservation of overall geometry with local deviations in incompletely reconstructed regions. Conclusions: Standard monocular microscope images combined with navigation-derived pose data can reconstruct millimeter-range 3D surfaces using a foundation-model-based pipeline. These results show technical feasibility in a controlled phantom setting and support further development toward objective quantification of operative exposure, image fusion, and characterization of working spaces for future surgical instrumentation. Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Medical Physics (physics.med-ph) Cite as: arXiv:2607.22773 [eess.IV] (or arXiv:2607.22773v1 [eess.IV] for this version) https://doi.org/10.48550/arXiv.2607.22773 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Thomas Bucher [view email] [v1] Fri, 24 Jul 2026 06:57:39 UTC (30,753 KB)

[CV-225] Generative Video Compression with Adaptive Score Distillation

链接: https://arxiv.org/abs/2607.22772
作者: Naifu Xue,Zhaoyang Jia,Haosen Li,Zihan Zheng,Jiahao Li,Bin Li,Xiaoyi Zhang,Qi Meng,Yuan Zhang,Yan Lu
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Diffusion models provide strong generative capabilities for video compression at ultra-low bitrates. Existing diffusion-based video codecs adapt base models originally developed for text-conditioned generation, whereas diffusion models designed and trained specifically for compression remain unexplored. To fill this gap, we introduce our Generative Video Codec (GenVC), built on a video diffusion model trained from scratch for compression. To our knowledge, this is the first compression-oriented video diffusion model. We realize this model directly in pixel space with a global-to-local hierarchy that recovers fine spatio-temporal details, enabling high-quality generative reconstruction from compressed representations. To accelerate inference, we distill the multi-step model into one step using distribution matching distillation (DMD). Applying DMD directly, however, drives the student toward motion-stalled reconstructions. We trace this to a teacher-side guidance failure: once student-induced perturbations leave the frozen teacher’s training region, its guidance can become misleading, causing DMD updates to reinforce rather than correct the student drift. To break the resulting feedback loop, we propose Adaptive Score Distillation, which gates DMD updates according to their alignment with the ground-truth direction, enabling high-quality reconstruction with coherent motion. Experimental results show that GenVC achieves state-of-the-art perceptual quality at ultra-low bitrates, with average bitrate savings of 62.5% at matched LPIPS and 71.3% at matched FID over GLVC. Unlike prior codecs that inherit billion-scale pretrained backbones, our diffusion model has only 478.0M parameters and decodes 1080p video in a single step at 15.1 fps on an A100 GPU.

[CV-226] Frequency-Aware Dual-Stream Learning for Balanced Realism and Fidelity in Electron Microscopy Imaging

链接: https://arxiv.org/abs/2607.22765
作者: Longmi Gao,Zhengkai Zhao,Pan Gao,Manoranjan Paul
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:Electron microscopy enables nanoscale cellular visualization but faces a trade-off between imaging resolution and acquisition speed. Existing learning-based methods rely on single-stream architectures that struggle to balance perceptual realism and quantitative fidelity, either over-smoothing details or generating unrealistic hallucinations. This work introduces a frequency-adaptive dual-stream architecture to resolve this conflict. Using discrete wavelet transform, we decompose images into low-frequency structures and high-frequency details, then employ a conditional diffusion model for realistic global synthesis and a transformer network for precise detail recovery. Experiments on the EMDiffuse dataset show the method achieves superior LPIPS and resolution ratio, substantially outperforming existing approaches. The method also shows strong generalization across diverse biological samples, supporting fast and reliable electron microscopy imaging for structural biology and nanotechnology applications. The source code and associated dataset are publicly available to facilitate further research.

[CV-227] pyALDIC: A Python Implementation of Augmented Lagrangian Digital Image Correlation with a GUI Adaptive Meshing and Mask-Aware Subset Splitting

链接: https://arxiv.org/abs/2607.22755
作者: Zixiang Tong,Jin Yang
类目: Image and Video Processing (eess.IV); Materials Science (cond-mat.mtrl-sci); Computer Vision and Pattern Recognition (cs.CV)
备注:

点击查看摘要

Abstract:pyALDIC is an open-source Python implementation of augmented Lagrangian digital image correlation (AL-DIC) for full-field displacement and strain measurement. The software combines a graphical user interface with a scriptable Python API and supports adaptive quadtree meshing, mask-aware subset splitting near cracks and holes, and selectable Local DIC and AL-DIC solver modes. Numba acceleration enables efficient analysis, while automated tests, documentation, and reproducible examples support reliable use acrossWindows, macOS, and Linux. Verification cases include synthetic displacement fields, rigid-body motion, Mode-I cracking, adaptive refinement, and experimental uniaxial tension. pyALDIC is distributed through PyPI, GitHub, and Zenodo under a BSD-3-Clause license for reproducibility. pyALDIC is openly available at this https URL.

人工智能

[AI-0] Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines

链接: https://arxiv.org/abs/2607.24692
作者: Jhonatan Tavori,Gur-Eyal Sela,Ion Stoica,Gil Zussman
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Distributed, Parallel, and Cluster Computing (cs.DC)
备注:

点击查看摘要

Abstract:Inference systems increasingly combine a fast path that returns predictions within the application’s latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be returned on time and combined with the fast path predictions. Across several application domains, we abstract this inference architecture as a fast path, a slow path, and a coordination layer with two functions: a router that invokes the slow path and a merger that decides whether to incorporate its returned predictions. In this work, we show that this new coordination layer exposes a new attack surface: shaped workload attacks, e.g., Yo-Yo bursts, can exploit contention at shared resources along the slow path to push benign users’ slow-path predictions past their latency deadlines. The merger then discards those predictions, while the fast path continues to return timely outputs. We refer to the resulting loss of slow-path accuracy benefits as accuracy collapse. We demonstrate accuracy collapse in a two-tier edge-cloud multi-object tracking pipeline in autonomous driving. In simulation, approximately 4,000 burst-shaped requests increase benign p99 latency from 92ms to 2s, nearly eliminating the benefit of the slow path’s cloud inference, reducing object tracking quality by 7.0 HOTA points on average. We further find that accuracy degradation can significantly vary (2.0-18.7 HOTA points), depending on the video intervals that are targeted in the attack, and that certain rare classes (e.g., stop signs) lose nearly half of their pre-attack prediction accuracy. These results show that workload attacks can degrade prediction quality without needing either access to model weights or victim data, and motivate research on attacks and defenses for routing, merging, scheduling, and resource isolation in these emerging inference pipeline architectures. Subjects: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Distributed, Parallel, and Cluster Computing (cs.DC) Cite as: arXiv:2607.24692 [cs.NI] (or arXiv:2607.24692v1 [cs.NI] for this version) https://doi.org/10.48550/arXiv.2607.24692 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-1] Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory and When Measuring Beats Accumulating

链接: https://arxiv.org/abs/2607.24667
作者: Maruthi Vemula,Neeraj Praneeth Gajula
类目: Artificial Intelligence (cs.AI)
备注: 8 pages, 3 figures, 3 tables

点击查看摘要

Abstract:A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will be reused, placing existing methods on one axis, the commit lag H : online filters and learned predictors commit at H=0 , while Belady’s offline optimum sits where the whole future is known. The missing regime in between, fixed-lag smoothing, waits a bounded number of steps, observes which items a correct near-future prediction attended to, and only then commits. This measurement, demonstrated utility, turns Belady’s unobservable future request into something we read off the model itself. We instantiate it as a training-free policy, RMM, a strict generalization of H2O that reduces to it exactly when the measurement is uniform. In controlled settings where reuse is endogenous and separated in time, demonstrated utility identifies used memory far better than accumulated attention, and a small bounded memory behaves like a much larger one. But on independent third-party benchmarks, run inside NVIDIA’s KVPress harness against its own SnapKV, H2O, and StreamingLLM implementations, the advantage mostly disappears: RMM is on par with H2O for single-turn question answering and loses to both H2O and SnapKV in a streaming multi-turn setting. The cause is simple: on natural text the model is correct about most tokens, so weighting attention by correctness barely changes it, and demonstrated utility collapses onto accumulated attention unless reuse is sharp and endogenous, which standard benchmarks do not exercise. Our contribution is the framework and an honest map of when measuring beats accumulating, not a new state of the art.

[AI-2] Reason -Mediated Behavioral Models for Auditing LLM Social Simulators

链接: https://arxiv.org/abs/2607.24649
作者: Atharva Pandey,Gautam Jajoo
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-derived reason pattern. We study this problem through a 94-person sunscreen concept test in which each respondent evaluated three product concepts and wrote open-ended rationales. We map those rationales into signed reason states Z , where positive signs support adoption and negative signs block it. This gives a practical audit: holding respondent descriptors D , category context K , and concept treatment X fixed, do human rationale-derived reasons help predict behavior Y , and can an LLM simulate the same reason state without seeing the human rationale or outcome? Human rationale-derived reasons substantially improve held-out prediction of purchase intent. LLM-simulated reasons are more brittle: they often sound plausible, but frequently echo the concept board rather than recover the respondent’s acceptance or rejection path. The paper contributes an evaluation framework for social simulators. Reason states do not identify natural causal effects by themselves, but they provide an interpretable test of whether a simulator’s stated reasons align with human evidence.

[AI-3] Efficiency Matters in Autonomous Research

链接: https://arxiv.org/abs/2607.24647
作者: Haiqian Yang,Yuan Cao
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of the solution-search process is an equally important but often overlooked dimension of performance. A strong AR system should not only produce high-quality results, but also reach them with as small a budget as possible. Search efficiency will become increasingly important as AR expands from domains with inexpensive verification, such as mathematics and coding, to real-world scientific settings in which solution evaluation may require costly physical experiments. To capture this dimension, we propose evaluating AR systems using the area under the curve (AUC) of the Pareto frontier, alongside final outcome quality. We compare several families of search algorithms, including hill climbing, beam search, tree search, and evolutionary search, across twelve systems-optimization tasks. We find that no single search structure is consistently the most efficient. We also show that search efficiency and final outcome quality are distinct performance dimensions: a method that eventually achieves the best result may nevertheless improve slowly and consume substantially more evaluation budget before reaching that result. Because the most effective search policy is generally unknown in advance, we introduce an adaptive procedure called fluid search, which uses a portfolio bandit to dynamically allocate a fixed evaluation budget across a forest of search processes. Across the evaluated tasks, fluid search achieves the highest overall search efficiency, closely matching the performance of a per-task oracle that is given the best search structure for each task in advance.

[AI-4] Agent ic Permissions Policy Algebra for Taint Confinement in LLM Agents

链接: https://arxiv.org/abs/2607.24625
作者: Arseny Kravchenko,Vadim Liventsev,Innokentii Konstantinov,Ildar Iskhakov,Matvey Kukuy
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: Preprint. Submitted to the 19th ACM Workshop on Artificial Intelligence and Security (AISec '26). 10 pages, 2 tables, 1 figure

点击查看摘要

Abstract:Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint tracking permanently taints an agent’s context upon reading unvetted data, severely restricting downstream utility. We present APPA (Agentic Permissions Policy Algebra), an IFC framework that resolves this usability bottleneck through engine-managed context branching and prospective acquisition enforcement. Before data acquisition occurs, APPA prospectively evaluates label descents and missing prerequisites, generating actionable remedy plans (Authorize, Accept). To inspect unvetted data without polluting the primary context, a label-seeded child trajectory is spawned, absorbing label descent locally and allowing a trusted sanitizer to return a bounded derivative to the unchanged parent. Governed by a two-monoid model over security labels and shared event logs, we formally prove parent label preservation and merge confinement. Finally, we evaluate APPA on a multi-turn tool-chaining benchmark across four models: it suppresses exfiltration (31%-50% down to 0%-7% attack success), and on three of the four, branching recovers a substantial share of the utility that taint tracking alone forfeits.

[AI-5] Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments Challenges and Future Directions

链接: https://arxiv.org/abs/2607.24589
作者: Zhimin Zhang,Chengzhen Ma,Jia Chai,Rongxin Zhan,Huansheng Ning,Lingfeng Mao,Dan Zhang,Suiping Jiang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The rise of Artificial Intelligence (AI) has significantly accelerated the global processes of digitization, informatization, and intelligence. Exploring how AI can leverage inherent characteristics to influence the development trajectory of IE is a topic that warrants further investigation. Given AI’s increasing prominence and role within IE, the paper analyzes this new form, examining both AI’s unique contributions to IE and its potential challenges. Firstly, the paper synthesizes the conceptual frameworks surrounding IE, decomposing them into manifestations in physical, social, and thinking spaces. Furthermore, the concept of Artificial Intelligence IE (AIIE) is introduced from a spatial perspective, with an exploration of the characteristics AI contributes to IE. Subsequently, the paper employs an evolutionary perspective to analyze the roles provided by AI during different development periods of AIIE. The paper then verifies the feasibility, effectiveness, and rationality of the AIIE’s definition and analyzes AIIE development from an evolutionary perspective using enterprise development examples. Finally, acknowledging AI’s inherent limitations, the paper examines potential challenges facing AIIE in the future from four perspectives, aiming to identify new research avenues for the further development of AIIE.

[AI-6] SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

链接: https://arxiv.org/abs/2607.24588
作者: Hang Ni,Weijia Zhang,Fan Liu,Mengqian Lu,Hao Liu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process. Although recent advances in Large Language Model (LLM) agents have enabled the automation of weather-related tasks, existing studies remain centered on isolated scientific tasks and overlook the chain of interdependent processes required for operational extreme-weather early warning. To bridge this gap, this study investigates automated end-to-end extreme-weather early warning through LLM agents. We first develop SIREN-Bench, a comprehensive benchmark comprising 600 question-answer instances across 19 tasks, and covering four individual warning procedures and an end-to-end warning chain. Evaluation on SIREN-Bench reveals substantial capability gaps in existing weather agent frameworks. This motivates us to develop SIREN, an experience-grounded agent framework inspired by experts’ use of historical cases, which combines an agentic execution environment integrating heterogeneous weather evidence and tools with a family of agent harnesses that exploit historical cases through retrieval, skill distillation, and predictive modeling. Extensive experiments demonstrate that SIREN outperforms weather-agent baselines on both individual warning procedures and end-to-end warning chains.

[AI-7] LLM -SoccerArena: Benchmarking LLM s on Real-World Predictions in Sports

链接: https://arxiv.org/abs/2607.24573
作者: Jonas Schröder,Jonas Schweisthal,Oliver Müller,Markus Weinmann,Stefan Feuerriegel
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (this https URL), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped, schema-validated forecasts of unresolved events, together with prompts, model versions, tool traces, and costs. The factorial design varies along four dimensions: (1) model version (e.g., GPT-5.5, Claude Opus 4.8); (2) information access; (3) prompting strategy, and (4) forecast horizon. We demonstrate LLM-SoccerArena through a large-scale evaluation of the 2026 FIFA World Cup, in which seven LLMs generated forecasts for all 104 matches and 15 tournament-related questions. We provide a detailed analysis of model performance across information access, prompting strategy, and forecast horizon. As a result, LLM-SoccerArena provides new evidence about the forecasting performance of state-of-the-art LLMs. For example, LLMs with web access outperform those without, but only by a small margin (i.e., a 0.023 improvement in Brier score). Overall, LLM-SoccerArena provides a flexible, open-source platform for prospective benchmarking of unresolved events. LLM-SoccerArena will be continuously updated, and can be directly applied to future national and international tournaments and league competitions.

[AI-8] RACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs

链接: https://arxiv.org/abs/2607.24563
作者: Federico Valletta,Giacomo Longo,Enrico Russo,Alessio Merlo
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:

点击查看摘要

Abstract:Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATTCK, yet extractor outputs remain fallible and are often stored without the evidence, provenance, and validation history needed to decide whether an individual mapping should be trusted. We present TRACE- CTI, a post-extraction claim-governance framework that preserves run-level Predictions, aggregates them into configuration-level GraphAssertions, materializes setup-deduplicated corroboration as ConsensusAssertions, and exposes only GraphAssertions backed by policy-compliant validation grounds. The framework retains native evidence granularity, complete extraction provenance, versioned trust decisions, and non-destructive revocation history. We evaluate TRACE-CTI on two public CTI corpora comprising 65 reports and 5,303 sentences, using a controlled 2 x 3 matrix of retrievers and generator families, incrementally ingested across six GraphVersions. All setups are incorporated without schema modification; provenance paths remain complete, operational scopes remain disjoint, and every trusted GraphAssertion has an active qualifying validation ground. Cross-generator-family setup pairs exhibit greater output diversity than same-family pairs. At the final graph state, increasing setup support from k = 1 to six-setup unanimity raises gold-aligned precision from 25.3% to 90.6%, while recall decreases from 88.2% to 16.3%. The graph also directly answers seven questions about provenance, trust, versioning, dependency, disagreement, and review-queue that the evaluated minimal flat output cannot fully answer without enrichment or reprocessing. These results support explicit, auditable governance of extracted TTP claims; the observed corroboration trajectory is descriptive and does not establish statistical independence or a causal model-family effect.

[AI-9] Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models

链接: https://arxiv.org/abs/2607.24562
作者: Murilo Salem,Luísa Böhm,Daniel Pontes,Anderson Ferrugem
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style. Conformal risk control (CRC) gives rigorous marginal risk guarantees for selective prediction with abstention, but marginal guarantees do not imply per-group ones: a model can meet the population budget while systematically over-exposing subgroups to errors. Under mild shift in group composition, standard CRC violates the budget in up to 47% of trials. We propose HG-CRC (Hierarchical Group-Conditional CRC), a post-hoc calibration framework enforcing simultaneous risk guarantees across all nodes of a user-defined group hierarchy. It applies a Bonferroni correction over nodes and a leaf-first policy that uses the most specific applicable threshold, falling back to coarser nodes when a finer one is uncertified or rejects the example. It needs only a held-out calibration set, with no retraining. We evaluate on three models (Qwen3-4B, Llama-3.1-8B-Instruct, Gemma-3-4B) and two benchmarks (ARC Challenge, MMLU-Pro) across eight configurations probing IID generalization, heterogeneity, mixture/domain/prompt/difficulty shift, label noise, and quantization. Main result: HG-CRC reaches an empirical 0% violation rate and WGER=0 on ARC Challenge for high-accuracy models (Qwen3-4B, Llama-3.1-8B). At 500 bootstrap trials these zeros are empirical upper bounds (true rate up to 0.6%), not certified. Results are benchmark-specific: on MMLU-Pro these models abstain entirely or (Llama) retain WGER=0.014. Gemma-3-4B, poorly calibrated here, degrades gracefully by abstaining. Participation cost vs. global CRC is 22 to 37 points. Ablations show hierarchical depth clears the budget: removing difficulty level returns violations to about 11%. Bonferroni is needed for the theoretical guarantee, though its empirical effect matters only with many nodes. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.24562 [cs.AI] (or arXiv:2607.24562v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.24562 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-10] BettiSplit: Topology-Guided Privacy-Aware Split Learning Against Feature Inversion and Gradient Leakage

链接: https://arxiv.org/abs/2607.24556
作者: Akarsh K.Nair,Muhammad Arifur Rahman,David Brown,Mufti Mahmud
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:

点击查看摘要

Abstract:Split learning enables collaborative model training by partitioning neural networks across clients and servers. However, improper split placement can lead to severe privacy leakage through intermediate representations. In this work, we propose a topology-guided framework for privacy-aware split learning based on the persistent Betti complexity of smashed activations. Through comprehensive layer-wise analysis, we show that privacy risk in split learning is highly non-uniform across layers and exhibits sharp transition regions that are not captured by architectural depth alone. In particular, feature inversion fidelity increases from negligible reconstruction to as high as 0.98 SSIM at deeper, privacy-critical split points. We further demonstrate that Betti complexity consistently identifies representation regimes associated with elevated feature-space privacy leakage across architectures and datasets. Leveraging this observation, we introduce BettiSafe, a topology-guided split selection strategy that identifies privacy-sensitive layers without requiring explicit attack execution. BettiSafe improves resistance to feature inversion by 2 to 5 times compared to depth-based heuristics while preserving classification accuracy. In addition, Betti-based regularisation increases inversion difficulty by nearly 5 x without degrading model utility, enabling a favourable privacy utility tradeoff. Overall, our results highlight topological complexity as a promising structural descriptor for secure, adaptive, and representation-aware split learning in real-world collaborative systems

[AI-11] LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding

链接: https://arxiv.org/abs/2607.24555
作者: Junsung Hwang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Attention keys are locally low-rank though globally high-rank: shared low-rank bases discard page-specific directions that a page’s own compact basis retains. LOCKS gives every page its own spectral summary (resident, about a tenth the cache’s size), reconstructs within-page logits, estimates each page’s attention mass by log-sum-exp, and attends only the top pages; selection itself reads no candidate keys or values. Selecting on this summary alone stays within about a point of the full cache on long-document QA (LongBench-v1), tracks the read-every-key oracle on retrieval-dense RULER down to the smallest budgets, and shows its largest margins on long-form reasoning (AIME26, MATH-500), where baseline selectors collapse. At its shipped 2048 -token budget LOCKS matches FullKV aggregate quality at 100 K + context while attending about 2% of the tokens, and halves per-token decode latency ( 2.0\times at 1 M tokens) against dense attention. LOCKS ships as a drop-in plugin for unmodified vLLM, with batched decode running in full CUDA graphs.

[AI-12] EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

链接: https://arxiv.org/abs/2607.24553
作者: Xiaocheng Fang,Jieyi Cai,Guangkun Nie,Haoyu Wang,Jiarui Jin,Yujie Xiao,Bo Liu,Chenyang He,Qinghao Zhao,Gaofeng Cheng,Hongyan Li,Shenda Hong
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings. Global ECG–text alignment may entangle modality-specific factors, while long-tailed finding distributions provide sparse positive supervision for low-prevalence conditions. We propose EchoBridge with Complementary Shared–Private Projection (CSPP) and Adaptive Prototype Boundary Calibration (APBC). CSPP maps each modality into shared and auxiliary private projections, reduces directional redundancy via within-modality orthogonality, and bidirectionally aligns normalized shared projections. APBC organizes the shared hypersphere with class-specific prototypes, training-frequency-adaptive angular margins, and spherical Riesz repulsion. We evaluate EchoBridge on EchoNext-Mini and independent PKUPH and SHTMU cohorts under four protocols: prompt-based inference without downstream classifier training, in-domain frozen linear probing, target-domain cross-center frozen linear probing, and source-only cross-center transfer, supplemented by finding-specific analyses. EchoBridge improves classifier-free AUROC, AUPRC, and F1 over the strongest baselines by 7.88, 5.61, and 4.54 points, respectively, and achieves the highest point estimates across all in-domain and target-domain probing budgets and both source-only transfer cohorts. Finding-specific analyses show gains for most conditions, including several low-prevalence valvular findings.

[AI-13] LLM -Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph

链接: https://arxiv.org/abs/2607.24551
作者: G{é}nesis Montenegro(WIMMICS),Mokhtar Boumedyen Billami,Catherine Faron(WIMMICS),Fabien Gandon(WIMMICS),Pierre Monnin(WIMMICS)
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate into operational systems. This paper presents a two-stage LLM-assisted workflow for French maintenance regulations: ontology engineering from a SEMLEG-based core ontology, followed by construction of an ontology-grounded French legal knowledge graph. The first stage consists in the open extraction of typed entities and triples from a stratified corpus sample, the normalization of labels through embedding-based fusion, and the induction of candidate object properties with their signature (domain and range). The second stage uses the resulting ontology to guide the closed extraction of triples and RDF graph construction over the full corpus. Experiments with GPT-4.1 and mistral-large-2512 show robust structured outputs, near-complete class alignment, and a substantial reduction of duplicated entities and predicates after fusion. Fewer than 20% of triples introduce unseen properties, while lower exact signature compliance reveals new domain-range combinations for existing predicates. These results point to predicate normalization and the validation of newly observed relation signatures as key refinement steps for industrial maintenance settings.

[AI-14] ask-Conditional Faithfulness Auditing of Multimodal LLM s for Grid Diagnosis

链接: https://arxiv.org/abs/2607.24539
作者: Tianqiao Zhao,Meng Yue,Jianhui Wang
类目: Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注:

点击查看摘要

Abstract:Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers task-specific evidence requirements and compares them with self-reported reliance and behavioral changes under controlled modality ablations. To resolve detected discrepancies, we design an evidence-gated correction and re-audit mechanism that regenerates failed responses under evidence constraints and independently re-ablates them to verify improved grounding without performance loss. Case studies evaluate three differently scaled LLMs on IEEE 39- and 118-bus scenarios. These results validate the framework ability to detect, diagnose, and correct task-conditional faithfulness failures.

[AI-15] Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls

链接: https://arxiv.org/abs/2607.24519
作者: Marzieh Zare
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE)
备注:

点击查看摘要

Abstract:Pretrained EEG foundation models are increasingly proposed for clinical decoding, but their transfer across populations and robustness to negative controls remain unclear. We benchmark six models (LaBraM, EEGMamba, CBraMod, REVE, BENDR, and BIOT) on five clinical tasks across four datasets using frozen linear probes with leave-one-subject-out, subject-grouped, or explicitly identified recording-level splits. Selected REVE findings are tested against random initialisation, random features, label permutation, scrambled-label fine-tuning, and projection sensitivity. On Korean dementia (CAUEEG, three-way), frozen REVE reaches 0.568 AUROC versus 0.769 for classical features; the ordering persists on a patient-disjoint held-out split (0.565 versus 0.768). Dataset identity is readily decoded from frozen embeddings (AUROC 1.000 at PCA-50; 0.9998 after band restriction and per-epoch z-scoring), whereas the same PCA-50 pipeline decodes Korean diagnosis at 0.528. A randomly initialised encoder also outperforms pretrained REVE on this task (0.659 versus 0.570). On Alzheimer’s disease, Gaussian random projection and PCA of the same pretrained embeddings perform similarly, and classical features nominally exceed REVE at the subject level. The clearest controlled positive is cross-subject ictal detection on CHB-MIT (n=23), where REVE achieves 0.793 AUROC, 9.2 percentage points above a randomly initialised encoder. These results show that EEG foundation-model conclusions depend strongly on evaluation unit, dataset shift, comparator strength, and targeted controls.

[AI-16] Making Mathematical Knowledge Explainable Accessible and Interoperable Through Large Language Model Integration ISWC

链接: https://arxiv.org/abs/2607.24512
作者: Jan Range,Björn Schembera,Dominik Göddeke
类目: Artificial Intelligence (cs.AI); Digital Libraries (cs.DL)
备注: Preprint submitted to 6th Wikidata Workshop@ISWC

点击查看摘要

Abstract:Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles. Knowledge bases such as the Mathematical Model Database (MathModDB) address this gap by providing curated, semantically rich representations of mathematical models. Built on Wikibase, the same open-source infrastructure underlying Wikidata, MathModDB utilizes Semantic Web technologies to support Linked Open Data, collaborative editing, and the storage of semantically enriched metadata, making it a domain-specific knowledge graph within the broader Wikidata ecosystem. However, access to MathModDB currently requires either navigating a complex web interface or proficiency in SPARQL and Wikibase APIs, posing significant barriers for potential users. In addition, the combination of such curated knowledge bases with actual research data stored, e.g., in Dataverse repository instances, remains a challenge. To overcome these limitations, we propose integrating Large Language Models (LLMs) with MathModDB via a Model Context Protocol (MCP) server that exposes a vector-indexed schema retrieval and Steiner-tree-based join planner, combining dialogue-based natural language interaction with curated, epistemically grounded knowledge. Although instantiated on MathModDB, the architecture can be applied to other Wikibase-based systems. We demonstrate that this approach enables epistemically grounded LLM usage, improves model explainability and accessibility beyond what the standard Wikibase interface offers, and simplifies interoperability with external databases and tools, such as Dataverse data repositories. We illustrate the benefits of combining the accessibility of an LLM with the epistemic safety of a curated knowledge base through the adaptability of the MCP protocol by two use cases involving mathematical models in the fields of continuum mechanics and enzyme kinetics.

[AI-17] UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

链接: https://arxiv.org/abs/2607.24507
作者: Xiaoyi Jiang,Jingyuan Li,Yixuan Jiang,Wei Liu,Yi Zhu,Zuoqiang Shi,Pipi Hu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling. However, adapting AR checkpoints across corruption kernels remains challenging because existing DLMs use different objectives and prediction parameterizations. We establish connections among SEDD, MDLM/GIDD, M2S, and Neural CTMC by expressing their conditional losses as a single generalized Kullback–Leibler objective over model reverse rates. We further derive conversions from clean-token predictions to concrete-score, posterior-mean, and exit-rate/jump parameterizations, yielding a shared (x_0) interface that supports switching between mask and uniform kernels. Building on these connections, we propose \ours, a simple continual pre-training approach for directly adapting pretrained GPT2 checkpoints to uniform-noise diffusion. Through systematic evaluation of 124M- and 355M-parameter models, we show that \ours steadily improves the trade-off between generative perplexity (GenPPL) and unigram entropy as the sampling budget increases from 16 to 256 steps. At 256 steps, \ours-S and \ours-M achieve GenPPL/entropy pairs of (97.783/5.2626) and (71.516/5.6669), respectively; no evaluated model at the same scale simultaneously outperforms \ours on both metrics. At both scales, \ours also achieves the highest WinoGrande, SIQA, and BBH accuracy among the compared diffusion models.

[AI-18] From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

链接: https://arxiv.org/abs/2607.24459
作者: Liwei Dong,Jiahao Zhao,Nan Xu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime experience into transferable procedural knowledge and persistent model improvement. This setting presents two challenges: trajectory-derived artifacts may encode source-specific repairs rather than cross-task computational mechanisms; and a weaker target model may be unable to operationalize an otherwise valid abstract procedure - an abstraction-execution gap. We introduce SciConsolidate, which contrasts verified successes and failures to induce cross-task procedures, selects them through a development-validation gate, and uses failure-informed, answer-free query synthesis to expand the consolidation data without requiring pre-existing reference answers. Because the target model may not directly execute these abstractions, a stronger model concretizes them into executable code supervision for standard, procedure-free SFT; a matched no-procedure teacher branch isolates the value of procedural guidance. On SciCode, runtime procedure injection improves Qwen3.6-27B by +3.85/+6.26 sub-step/main-problem points, but yields almost no aggregate main-problem gain for Qwen3.5-9B, providing operational evidence of the abstraction-execution gap. After procedure-guided concretization, the 9B student improves under procedure-free deployment by +3.89/+6.25 points over the no-procedure SFT control and by +5.62/+11.25 over the original 9B model. These results establish an experience-to-capability pathway for scientific computing and provide a practical starting point for scaling self-improving scientific assistance.

[AI-19] DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

链接: https://arxiv.org/abs/2607.24434
作者: Dengke Han
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 9 pages, 6 figures

点击查看摘要

Abstract:Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU. In this setting, self-speculative decoding faces a new bottleneck: increasing the draft expert set improves accuracy but triggers extra expert loading, while cheap small-footprint drafts have low acceptance; moreover, verifying a multi-token block activates the union of target experts and is no longer close to one target step. We propose DraftExpert, an expansion-aware self-speculative decoding framework for expert-offloaded MoE inference. DraftExpert trains one lightweight accelerator-resident draft expert per layer by self-distilling residual, logit/token, and router-agreement signals from the frozen target MoE. At inference time, it uses a fixed-footprint shared+top-1+draft-expert drafter together with confidence–expansion truncation and target-expert prefetching, while final tokens are still exactly verified by the target model. On DeepSeek-V2-Lite and Moonlight-16B-A3B across CPU-GPU and Flash-NPU offload, DraftExpert improves decode throughput by 1.45x on average, raises draft acceptance to 84~87%, and achieves 86~88% prefetch hit rates.

[AI-20] Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

链接: https://arxiv.org/abs/2607.24419
作者: Jinliang Deng,Yiming Niu,Yibo Pan,Zhiqi Shao,Qin Luo,Yongxin Tong
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and iteratively revise classifier designs. Recent LLM-based agents have demonstrated the potential for automated model design, but when guided only by aggregate performance metrics, they lack insight into why individual cases fail and how the classifier should be revised. We present RecursiveECG, an evidence-driven LLM-as-Designer framework in which an LLM serves as an offline model designer that refines ECG classifiers based on concrete failures and objective ECG evidence. To ground failure diagnosis in executable evidence, Criteria-to-Measurement Compilation converts curated ECG criteria into validated deterministic functions that produce reproducible, reference-backed measurements for individual ECGs. Building on these measurements, Evidence-Grounded Failure Review analyzes failed and comparator cases by jointly considering raw waveforms, measurements, and model outputs, enabling the LLM to diagnose classifier limitations and formulate targeted revisions. Candidate revisions are executed and re-evaluated under a fixed problem contract, and only evidence-supported updates are retained. The resulting predictor is frozen after refinement and requires no LLM inference during deployment, while an audit trail links each accepted revision to its supporting evidence. Across PTB-XL, Georgia, and CPSC2018, RecursiveECG consistently outperforms strong baselines, achieving an average relative improvement of 10.0%. Extensive ablation and transfer studies further validate the effectiveness of its evidence-grounded refinement process.

[AI-21] he SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing

链接: https://arxiv.org/abs/2607.24396
作者: Stefan Scholze,Johannes Partzsch,Sebastian Höppner,Florian Kelber,Andreas Dixius,Marco Stolba,Sirine Arfa,Marc Berthel,Georg Ellguth,Jim Garside,Hector A. Gonzalez,Stephan Hartmann,Thomas Kiel-Hocker,Dongwei Hu,Matthias Jobst,Khaleelulla Khan Nazeer,Tim Langer,Chen Liu,Gengting Liu,Matthias Lohrmann,Mantas Mikaitis,Felix Neumärker,Amirhossein Rostami,Stefan Schiefer,Tilo Schubert,Delong Shang,Bernhard Vogginger,Yexin Yan,Steve Furber,Christian Mayr
类目: Emerging Technologies (cs.ET); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 19 pages, 13 figures

点击查看摘要

Abstract:In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphic hardware has long been advocated as an upcoming alternative to deep networks, taking inspiration from the brain for achieving unprecedented energy efficiency. However, demonstrations of these gains only recently began to grow in complexity and real-world applicability. With SpiNNaker2, we present a chip that bridges the gap between deep networks and neuromorphic computing and allows for flexible exploration of computing approaches that combine both worlds. It features 152 processing elements equipped with an ARM M4F processor and dedicated accelerators, an extended SpiNNaker routing fabric for scalable event-based communication and a range of external interfaces for system integration, including Gbit Ethernet and an LPDDR4 memory interface. We demonstrate performance and efficiency of the SpiNNaker2 chip for neuromorphic and deep network workloads, as well as novel event-based computing approaches. For deep network workloads, the chip achieves up to 4.5 TOPS in high performance mode and up to 2.7 TOPS/W efficiency in high efficiency mode for INT8 workloads. The chip supports spiking neural networks with 150000 neurons and 1.8 billion synaptic events/s when simulated with a 1 ms time step. Its low baseline power of less than 250 mW allows for efficiency even under varying workload conditions, allowing to explore sparse and event-based modes of computation. All this demonstrates the chip’s capabilities as a universal hardware platform for scalable brain-inspired computing and its combinations with mainstream deep network approaches.

[AI-22] Regulating for AI Legitimacy

链接: https://arxiv.org/abs/2607.24391
作者: Gilad Abiri
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:AI systems already govern. They rank speech and allocate attention, filter applicants and triage claims. The dominant frame for AI governance, alignment, asks whether such systems pursue the right objectives safely. It cannot answer a prior question: by what right are those objectives set and enforced? This Article argues that legitimacy is an autonomous regulatory objective, distinct from alignment and not secured by it. Legitimacy here is sociological: the belief among those subject to power that it is exercised rightfully. Performance does not produce that belief. We already have the proof of concept. Social media and search delivered enormous gains on every familiar metric and still triggered a legitimacy crisis, because publics questioned who authorized a handful of firms to set the rules of speech, visibility, and knowledge. It is possible to build a benevolent AI and still face a political crisis over its authority. The Article maps three sites where AI legitimacy falters: opacity, which blocks audiences from forming justified beliefs; private power, where firms exercise public-facing authority without recognizable authorization; and administrative automation, which strains reason-giving, participation, and review inside the state. It then asks what law can contribute. Thin legality (publicity, stability, consistent application) signals non-arbitrariness and buys real recognition, but invites legitimacy-washing when form drifts from practice. Thick legality supplies what form cannot: public authorship of the rules that bind. Three portable principles follow. Integration seats consequential AI rule-setting in venues a polity already treats as authoritative. Familiarity presents rules and reasons in locally credible forms. Contestation guarantees a credible second look with real remedies. Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.24391 [cs.CY] (or arXiv:2607.24391v1 [cs.CY] for this version) https://doi.org/10.48550/arXiv.2607.24391 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Gilad Abiri [view email] [v1] Mon, 27 Jul 2026 13:06:38 UTC (222 KB)

[AI-23] Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization

链接: https://arxiv.org/abs/2607.24354
作者: Haoyue Liu,Xiaoyu Ma,Ye Chen,Yuexian Zou,Xiaoying Tang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However, on multimodal tasks, the effectiveness of APO is fundamentally bottlenecked by a blind feedback channel: the optimizer reads the question, the prediction, and the gold answer, but never the input image on which the model failed, and therefore cannot diagnose visually grounded errors. As a remedy, we introduce Cross-Modal Visual Feedback (CMVF). CMVF incorporates (1) a failure-conditioned visual diagnosis stage, in which a stronger optimizer VLM inspects each failed image without access to predictions or labels, and (2) an error-aware aggregation stage that compresses these observations into reusable, task-level visual blind-spot patterns that drive the prompt rewrite. Crucially, the image is consumed only during optimization; the deployed artifact is an ordinary text prompt that runs at the same inference cost as any text-only baseline. Extensive results across 12 VQA datasets and 4 target VLMs demonstrate that CMVF consistently ranks first, improving over the strongest baseline on every target by 2.4 points on average, with gains of up to 6.5 points on individual benchmarks. Moreover, the optimizer self-organizes into expert-style visual checklists that transfer across models without re-optimization.

[AI-24] DeepFaith: Evidence-Grounded LLM s for Faithful Incident Reporting in Multi-Stage APT Defense

链接: https://arxiv.org/abs/2607.24348
作者: Trung V. Phan,Tri Gia Nguyen,Thomas Bauschert
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: This paper has been submitted to the IEEE International Conference on Network and Service Management (CNSM) 2026

点击查看摘要

Abstract:Advanced Persistent Threats (APTs) are difficult to detect and interpret due to their multi-stage and stealthy nature. While recent autonomous defense systems leverage provenance graphs and learning-based models for detection and mitigation, their outputs remain largely machine-oriented and difficult for analysts to interpret. Large language models (LLMs) offer a promising interface for report generation, but often produce hallucinated or weakly grounded content. In this paper, we propose DeepFaith, an evidence-grounded framework for faithful incident reporting in multi-stage APT defense. DeepFaith transforms structured outputs from autonomous defense and explainability modules into natural-language reports that are explicitly aligned with underlying system evidence. The framework integrates a unified evidence representation, evidence-grounded prompting, faithfulness-aware generation, and post-generation verification to ensure that all generated statements are supported. Experiments in a realistic enterprise testbed demonstrate that DeepFaith improves faithfulness from 0.68 to 0.92, reduces unsupported claims from 0.32 to 0.08, and increases temporal consistency from 0.6 to 0.88, while maintaining concise reports and lower error rates than existing template-based and LLM-based solutions. These results show that evidence-grounded generation enables reliable, interpretable, and actionable reporting for security operations centers.

[AI-25] Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age

链接: https://arxiv.org/abs/2607.24341
作者: Weijie Xia,Stefanie Horian,Hanyue Huang,Queena K. Qian,Jie Yang,Pedro P. Vergara Barrios
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent studies use Large language models (LLMs) to simulate human opinions and decisions by prompting models with demographic, attitudinal, or persona-based descriptions. Yet such simulations rarely model the practical, cognitive, or social frictions that shape how people respond to policy interventions. Perceived transaction cost (PTC) provides a useful lens for modeling the practical frictions that shape policy responses, such as information burden, administrative effort, coordination demands, and perceived uncertainty. We use this lens to develop a friction-aware persona modeling approach for LLM-based simulation. In the context of energy-efficient renovation (EER), tenants are represented not only by who they are demographically, but by how they perceive the costs, benefits, barriers, and uncertainties associated with proposed renovation plans. Using survey data collected from 1,068 citizens in the Netherlands, comprising approximately 40,548 survey question and answer pairs, we compare prompt-only and fine-tuned settings across GPT-3.5-turbo, Ministral-8B-Instruct, and Llama-3.1-8B-Instruct, and evaluate supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO) for local open-weight models. Results show that incorporating PTC-based personas and reasoning consistently improves model performance across both prompt-only and fine-tuned settings, suggesting that PTC-based persona design provides a useful bridge between institutional policy theory and interpretable LLM-based policy simulation. Code is available at this https URL.

[AI-26] Unequal Trips Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination

链接: https://arxiv.org/abs/2607.24336
作者: Nicole Hu,Mingtao Zhang,Haoyang LI,Chen Jason Zhang,Li Qing
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed across trips and regions. We conduct a distributional audit on three real-city road-network and taxi-demand datasets from Manhattan, Chicago, and San Francisco. The audit reveals pervasive trip-length inequity whose direction depends on the city and coordinator. After accounting for trip length, spatial inequity becomes more pronounced as demand grows and is consistently stronger when trips are grouped by origin rather than destination. These findings motivate SPatially Aware RErouting (SPARE), a budgeted online coordination framework that assigns limited replanning capacity to delayed vehicles and redirects them using recently observed waiting pressure. SPARE provides a per-review decision guarantee and explicitly bounds online route updates. Experiments on all three datasets against six representative baselines show that SPARE delivers the strongest joint efficiency-fairness performance while retaining city-scale scalability. The results demonstrate that bounded congestion-responsive rerouting improves performance and equity without full-fleet replanning.

[AI-27] From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agent ic Search

链接: https://arxiv.org/abs/2607.24280
作者: Junlin Liu,Jiangwang Chen,Zixin Song,Shuaiyu Zhou,Chunji Lv,Hank Wu,Kailin Jiang,Jinyang Wu,Bohan Yu,Chenxi Zhou
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this supervisory signal, conventional logit-matching is precluded by hidden logits and mismatched tokenizers, whereas raw natural language trajectory imitation transfers superficial stylistic artifacts rather than core reasoning competence. To address the heterogeneous distillation problem and bridge the distribution gap, we propose Multi-Agent Protocol Distillation (MAPD), a joint distillation and RL framework uses a structured, style-normalized protocol as an intermediate representation. An offline multi-agent system (MAS) decomposes each query, retrieves supporting evidence, repairs failed searches, and converts the resulting exploration trace into a JSON protocol containing the task type, reasoning plan, and extractive grounding facts. During training, the protocol is provided only to a privileged branch of the student policy, whose token distributions furnish a dense distillation signal alongside the sparse RL objective. Extensive evaluations across seven QA benchmarks demonstrate that MAPD consistently outperforms competitive distillation and RL, achieving average success rates of 39.4% on Qwen3-1.7B and 44.4% on Qwen3-4B. Crucially, the framework generalizes robustly across diverse proprietary teachers while effectively mitigating the student policy from style drift and verbosity degeneration.

[AI-28] A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health

链接: https://arxiv.org/abs/2607.24275
作者: Oluwadara Adedeji,Michael Mayowa Farayola,Jeff Brozena,Irina Tal,Regina Connolly,Mark Matthews
类目: Logic in Computer Science (cs.LO); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: Accepted at the CIBB 2026 conference ( this https URL )

点击查看摘要

Abstract:Ethical governance of AI-driven systems is often expressed through high-level principles and static documentation, creating a gap between regulatory requirements and system-level verification. This challenge is particularly acute in digital phenotyping, where continuous behavioural data raises concerns around consent, privacy, and fairness. In this paper, we propose a computational ethical framework for AI-driven digital phenotyping system in which ethical requirements are formalised as deontic temporal logic constraints, alongside a conceptual ethical agent that oversees the system and ensures that any supervised system satisfies the specified constraints. Using a case study involving financial data and mental health, we model key ethical properties and verify them using the Z3 Satisfiability Modulo Theories (SMT) solver. Our evaluation shows that the framework is logically consistent and that violations of the specified ethical properties are ruled out within the formal model through counterexample-based verification. This presents early research enabling continuous, machine-verifiable ethical checking, moving beyond retrospective compliance based on static documentation. We discuss limitations, including the need for real-world verification with data, the challenge with subjectivity and contextual sensitivity, the need for human oversight, and outline how such approaches can support the development of digital phenotyping and AI systems with continuous and auditable ethical guarantees.

[AI-29] Physics-Guided Generative AI for Property-Targeted 3D Porous Media Design

链接: https://arxiv.org/abs/2607.24274
作者: Peng Wang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 15 pages, 8 figures

点击查看摘要

Abstract:Inverse design of three-dimensional porous media is central to applications in filtration, catalysis, energy storage, fuel cells, thermal management, and biomedical scaffolds, but remains challenging because many distinct pore geometries can share similar porosity or permeability while small structural changes can strongly affect transport behaviour. This paper proposes a physics-guided generative AI framework for property-targeted porous media design, combining a property-aware variational autoencoder, a conditional latent diffusion model, and an independently trained differentiable structure-to-property surrogate. The framework learns a compact, physically informative latent design space, generates porous structures conditioned on target porosity and directional permeability, and refines generated samples using property-level feedback during denoising and decoding. Experiments on procedurally generated structures and real micro-CT porous-media datasets show improved target-property matching, directional permeability control, and property correlation compared with representative property-aware variational-autoencoder and latent-diffusion baselines. The results demonstrate a scalable route towards controllable inverse design of complex porous geometries and establish a foundation for simulation-informed generative AI tools in engineering and advanced materials discovery.

[AI-30] Generative Artificial Intelligence (GenAI) to convert images of queuing networks into verifiable simulation models: an open-weight LLM workflow approach

链接: https://arxiv.org/abs/2607.24259
作者: Thomas Monks,Alison Harper,Amy Heather,Navonil Mustafee
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent work has explored the use of Large Language Models (LLMs) to automate simulation model building, typically by generating executable code directly from natural language descriptions. However, this raises challenges for verification and reproducibility particularly for users without programming expertise. We propose Sketch2DES, a sketch-to-simulation workflow that converts diagrammatic representations of queuing networks into verifiable discrete-event simulation models using open-weight LLMs. The workflow has three stages: (1) translation of a diagram into a semi-structured textual description using a multimodal LLM; (2) conversion into schema-validated structured data (JSON) via an LLM with a reflection-based verification loop; and (3) deterministic transformation into an executable simulation model using a software adapter. Intermediate artefacts can therefore be inspected and automatically validated before execution. We evaluate the approach on eight queuing-network diagrams of varying complexity. The workflow achieved high reliability for all stages, and results were statistically indistinguishable from human-coded and analytical benchmarks. Compared to direct code generation, the workflow improves reproducibility, transparency, and verifiability, while reducing the need for programming expertise. Limitations include restricted model scope and dependence on accurate visual interpretation. The results demonstrate the feasibility of structured, workflow-based model generation as a robust foundation for LLM-assisted simulation modelling.

[AI-31] ML-based Predictive Models for Power Consumption in Virtualised O-RANs

链接: https://arxiv.org/abs/2607.24256
作者: Rishu Raj,Genevieve Akude,Urooj Tariq,Daniel Kilper
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As communication networks adopt virtualized and disaggregated architectures, achieving energy efficiency has become increasingly important for both economic and environmental reasons. Traditional methods for power modeling are inadequate in these dynamic software-defined environments due to their inability to model complex and nonlinear factors affecting energy use. We investigate the use of feature extraction and regressor-based machine learning methods for predicting power consumption in virtualized open radio access networks (O-RANs), utilizing datasets from a hardware-instrumented testbed. We test three variants of deep neural networks (DNNs), namely, a standard DNN, a regularized DNN, and a hybrid model combining DNN-based feature extraction with an XGBoost regressor. We evaluate the performance of these models for various system parameters such as transmission gain, modulation/coding schemes, and airtime. We show that the hybrid model consistently outperformed others, achieving a mean relative error below 0.5%. Results suggest hybrid models like DNN-XGBoost offer superior accuracy and could be integrated into O-RAN management tools to enable more energy-efficient network orchestration in future networks.

[AI-32] Epistemic Norms for AI Safety and Alignment Research

链接: https://arxiv.org/abs/2607.24243
作者: Keivan Navaie
类目: Artificial Intelligence (cs.AI)
备注: 36 pages

点击查看摘要

Abstract:Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes—\it capability profile, demonstrating the absence of hazardous behaviours rather than the presence of positive capabilities, and \it risk profile, bounding worst-case outcomes under fat-tailed uncertainty rather than optimising average-case performance—and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps, we propose \sc ECAISA, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an information-hazard adjudication procedure, and seven anti-gaming mechanisms. \sc ECAISA does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than certification as its governance target.

[AI-33] Every Client Is an Environment: Federated De-confounding for Spatio-Temporal Forecasting

链接: https://arxiv.org/abs/2607.24218
作者: Qingxiang Liu,Anqi Liang,Heng Wang,Yuxuan Liang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Federated learning has emerged as a promising paradigm for spatio-temporal forecasting (STF), enabling collaborative model training without sharing raw observations. Existing federated STF methods primarily regard cross-client heterogeneity as an optimization challenge and mitigate it through personalized approaches. However, such heterogeneity fundamentally stems from diverse \emphenvironmental conditions, and these methods capture environment-specific forecasting patterns, hardly generalizing under environmental shifts. Our key insight is that the environmental diversity across federated clients should be exploited, as they provide \emphcomplementary observations of the same underlying spatio-temporal system. Based on this insight, we propose \method, a novel federated de-confounding framework that \textbftreats clients as distinct causal environments. \method leverages the client heterogeneity as distributed environmental evidence and learns a global prototype codebook to capture shared environmental regimes. We further derive a theoretical federated de-confounding bound that is linearly controlled by the averaged confounding strength. Extensive experiments demonstrate that \method consistently outperforms federated baselines, while providing transferable, interpretable, and communication-efficient environmental representations.

[AI-34] Myopia Prevention and Control 3.0: Artificial Intelligence–Driven Risk Stratification Proactive Monitoring and Personalized Intervention

链接: https://arxiv.org/abs/2607.24187
作者: Tieniu Wang,Cangzhu Huang,Qianhui Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The convergence of artificial intelligence (AI), digital sensing, and ubiquitous computing has created an unprecedented opportunity to transform myopia prevention from a reactive, population-based model into a proactive, precision-driven one. Despite evidence that half the world’s population will be myopic by 2050, conventional approaches—school-based vision screening (Phase 1.0) and evidence-based risk factor management (Phase 2.0)—have proven insufficient. We review the emergence of Myopia Prevention and Control 3.0, defined by AI integration across three interconnected domains forming a closed-loop pipeline: (1) AI-driven risk stratification predicting individual-level risk through machine learning on multimodal data; (2) AI-enabled proactive monitoring via wearables, smartphones, and school screening networks; and (3) AI-powered personalized intervention with closed-loop feedback. We critically evaluate evidence across each stage, discuss challenges in data quality, model validation, ethics, and equity, and outline future directions including multimodal foundation models, digital twins, and causal machine learning.

[AI-35] Monitoring Post-Disaster Urban Recovery Using High-Resolution SAR Time Series and Unsupervised Learning: Evidence from the 2023 Türkiye-Syria Earthquake

链接: https://arxiv.org/abs/2607.24180
作者: Luigi Russo,Deodato Tapete,Silvia Liberata Ullo,Paolo Gamba
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Submitted to IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS)

点击查看摘要

Abstract:Monitoring post-disaster recovery is essential for understanding how urban systems rebuild and progressively return to functionality. However, tracking reconstruction remains difficult because reliable ground-truth information is often scarce and recovery processes evolve over time. This paper proposes an unsupervised framework for recovery monitoring based on multi-temporal synthetic aperture radar (SAR) observations and deep-learning anomaly detection. COSMO-SkyMed time series are used to identify persistent temporal anomalies associated with reconstruction activities and to generate spatially explicit recovery maps. The framework is applied to four cities severely affected by the 2023 Turkiye-Syria earthquakes, revealing heterogeneous reconstruction dynamics across different urban contexts. The results show spatially structured patterns of persistent anomalies related to reconstruction over damaged and cleared areas, temporary container settlements, and new residential districts. Comparison with nighttime-light recovery indicators derived from SDGSAT-1 data highlights the complementary nature of the two modalities: nighttime lights reflect the restoration of electricity supply and nighttime socioeconomic activity, whereas SAR anomalies capture structural changes in the built environment and may reveal reconstruction at earlier stages. The results demonstrate that multi-temporal SAR data combined with unsupervised learning provide an effective and scalable approach for monitoring post-disaster reconstruction when labeled recovery datasets are unavailable.

[AI-36] Falsifiable Commitment Planning for Self-Correcting Web Agents

链接: https://arxiv.org/abs/2607.24167
作者: Guangyi Liu,Huan Zhao,Quanming Yao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, or plan assumption no longer supports the user instruction. Existing agents can plan, reflect, or reuse experience, but their plans rarely specify the evidence under which an active step should still be trusted. We propose FCPAgent, a falsifiable commitment planning framework for robust long-horizon web agents. FCPAgent represents each plan step as a Falsifiable Commitment Unit (FCU): a subgoal grounded in a reusable skill, together with confirming evidence, falsifying evidence, and a confidence score. Execution is organized as a plan-test-repair loop. The hybrid commitment testing module checks candidate actions before they modify the browser and checks observations after execution; for efficiency, it combines lightweight evidence matching with LLM-based diagnostic verification. When evidence falsifies a commitment, scope-aware repair localizes the contradiction to the execution, skill, or planning level and revises the smallest adequate part. On WebArena, FCPAgent achieves a 13.8% relative improvement in average success over the strongest baseline, with especially large gains on long-horizon tasks.

[AI-37] Agent -UCT: Upper Confidence Bounds Applied to Trees for Agent ic Workflow Optimization with Cost-Awareness

链接: https://arxiv.org/abs/2607.24162
作者: Yang Li,Hai Liu,Dian Shao,Yu Wang,Xiyu Chen,Sergey Volkov,Bozhi Wang,Ziyu Sun,Sihang Liu,Ye Luo,Xiaowei Zhang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box optimization, and standard tree search methods - do not explicitly exploit the compositional structure of these workflows, leading to redundant computation and inefficient budget allocation. We introduce Agent-UCT (Agent-based Cost-Aware Upper Confidence Bounds Applied to Trees), a tree search algorithm that extends UCT with a reuse-aware regularization term derived from a bipartite prefix reuse graph. Agent-UCT biases selection toward branches that leverage previously materialized configuration prefixes, reducing redundant execution while maintaining effective exploration. Our framework, RAGSpace, unifies heterogeneous RAG components from LongRAG, LightRAG, and Self-RAG into a five-dimensional configuration space, enabling systematic cross-framework recombination. WTB (Workflow Test Bench) provides deterministic replay, content-addressable caching, and transactional consistency, ensuring that intermediate states are materialized once and reused across the search. Experiments on HotpotQA and UltraDomain demonstrate that Agent-UCT identifies configurations with the highest out-of-sample performance among the evaluated fixed framework presets. Under full-pool evaluation, bipartite prefix reuse reduces logical search cost by 73.6% relative to the no-prefix-sharing cost upper bound. Compared with full-pool evaluation, sampling-based evaluation further achieves a 4.2x wall-clock speedup. Agent-UCT, RAGSpace, and WTB together provide a unified framework for cost-aware, reproducible, and compositionally efficient agentic workflow optimization.

[AI-38] A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

链接: https://arxiv.org/abs/2607.24148
作者: Zhuoran Song,Haozhe Jiang,Chunyu Qi,Minnan Pei,Gang Li,Xiaoyao Liang,Haibing Guan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-precision workloads, leaving substantial redundancy in both memory and computation underexploited. In this paper, we propose VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics. We first introduce MotionVQ, a motion-aware vector quantization scheme that dynamically adjusts quantization precision based on the robot’s execution state, reducing memory access while preserving task success rate. We then propose a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation, eliminating redundant multiplications through spatial aggregation and temporal reuse of centroids. To realize these optimizations, we design an accelerator that efficiently supports dynamic precision selection and centroid-reuse computation. Experimental results show that VQVLA achieves 6.5x, 2.8x, 1.9x, 3.3x, and 4.3x speedup over the A100 GPU, Dadu-Corki, LUT-DLA, CodeGEMM, and ShiftAddLLM, respectively, with negligible accuracy degradation. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.24148 [cs.AI] (or arXiv:2607.24148v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.24148 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-39] Scaling GUI Agents with Visual State Transitions

链接: https://arxiv.org/abs/2607.24112
作者: Xiangyan Liu,Kaixin Li,Haonan Wang,Biao Wu,Meng Fang,Longxu Dou,Chao Du,Michael Qizhe Shieh,Tianyu Pang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predicting actions from state changes) and forward dynamics (predicting next states from current states and actions). This optimization equips the model with better action-grounded visual representations and an internal world model of GUI dynamics. When subsequently fine-tuned on trajectories with task instructions, our STP-trained models consistently outperform baselines trained solely via direct trajectory fine-tuning across agent benchmarks in both desktop and mobile GUI scenarios (AgentNetBench, AndroidControl, and GUIOdyssey). Further empirical studies show that joint dynamics optimization yields stable improvements over single-objective training, and downstream performance scales steadily with the volume of transition data.

[AI-40] MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents

链接: https://arxiv.org/abs/2607.24097
作者: Yiwen Ma,Songjun Tu,Qichao Zhang,Dong Li,Linjing Li,Dongbin Zhao
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This retrieval-as-evidence paradigm assumes retrieved memories are already suitable for reasoning, leaving the answer model to resolve redundancy, conflicts, and weak relevance while incurring substantial context overhead in long-term memory tasks. We propose MemChain, a trainable post-retrieval memory policy that transforms retrieved candidates into answer-facing active memory, represented as a compact and grounded evidence context. Given a user query and retrieved candidates, MemChain first generates a question-conditioned evidence plan, then constructs an ordered grounded evidence trace that organizes retrieved memories according to their semantic roles and dependencies, and finally executes explicit memory actions to produce a concise evidence context for answer generation. To train the mediator, we introduce a two-stage learning framework. Supervised trace learning first teaches the policy to generate structurally valid plans, traces, actions, and evidence contexts. We then propose Trace-Guided Memory Policy Optimization (TMPO), a reinforcement learning objective that optimizes the memory policy using downstream answer quality while jointly encouraging trace grounding, evidence support, structural validity, and answer stability across multiple rollouts. Experiments on LoCoMo and LongMemEval-S demonstrate that MemChain consistently achieves state-of-the-art performance across both closed-source and open-weight frozen answer models while substantially reducing the memory context passed to the answer model.

[AI-41] owards High-Level Semantic Intelligence

链接: https://arxiv.org/abs/2607.24082
作者: Xiujie Song,Gefei Yang,Yining You,Jiahui Gan,Qi Jia,Shota Watanabe,Tianxi Wan,Mengyue Wu,Kai Yu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear trajectory from simple to complex semantic processing. While early AI systems mainly addressed tasks involving direct and literal semantic perception or expression, contemporary systems are increasingly expected to perform more sophisticated cognitive reasoning, enabling the understanding and generation of High-Level Semantics (HLS). A similar trajectory can also be observed in human cognitive development. We define this transition as the shift from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI). However, this issue has not yet been systematically and comprehensively examined in prior work. Motivated by this gap, this survey reviews the development of AI semantic intelligence from the perspective of semantic complexity. We systematically survey existing research on HLS tasks, including humor, sarcasm, metaphor, empathy, persuasion, narrative, and other general HLS phenomena, across text, speech, vision, and multimodal scenarios. Specifically, we summarize data construction methods, modeling and optimization strategies, and evaluation methodologies for both understanding and generation. HLS is essential for advancing AI toward genuinely human-like intelligence. By synthesizing existing methods and insights from the perspective of semantic intelligence, this survey aims to support the continued development of AI toward HLSI.

[AI-42] MiSS: A Logic-Driven Explanation of Minimal Sufficient Coalitions for Point Cloud Classifiers

链接: https://arxiv.org/abs/2607.24074
作者: Mengda Xing(UA, CRIL),Jean-Marie Lagniez(UA, CRIL)
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We present MiSS, a black-box, query-based framework for explaining 3D point cloud classifiers through perturbation-relative sufficiency reasoning. MiSS treats a superpoint partition as an interpretable abstraction layer and asks whether the original prediction can be certified from a minimal coalition of geometric regions under a specified perturbation distribution. Unlike abductive explainers that require Boolean feature spaces or white-box logical encodings of the predictor, MiSS separates candidate proposal from verification: a weighted MaxSAT procedure proposes coalitions using a heuristic adaptive cardinality floor, certified exact-size fallback, a safely tightened upper bound, blocking clauses, and a surrogate acquisition heuristic learned from previous oracle evaluations, while a blackbox statistical oracle decides sufficiency from prediction queries. The system returns a statistically verified sufficient coalition as a binary attribution, with minimum cardinality guaranteed when certified search completes. Experiments on ModelNet40 and ShapeNet with PointNet and PointMLP classifiers show higher precision and coverage than rule-based baselines in most settings, with lower explanation time than exhaustive search.

[AI-43] he Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

链接: https://arxiv.org/abs/2607.24063
作者: Keyu Li,Jin Gao,Dequan Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that factuality costs. Static leaderboards score factuality in isolation and treat compute as free, so they cannot tell a genuinely better system apart from one that simply spends more. Consider a ranking reversal. A brute-force Best-of-4 agent posts the higher raw factuality score (H-Score 0.9169 vs 0.9103) and would top a static leaderboard, but once cost is counted it is the worse system, losing on Q-Score (0.5169 vs 0.5217) at roughly four times the tokens and latency, under a reported cost weight whose sensitivity we sweep. So the system that tops a static leaderboard can be the worse one to deploy. To make this trade-off visible, we introduce MAS-HQ (Multi-Agent System Hallucination Quest), a resource-aware evaluation protocol. It wraps any factuality detector and normalizes for cost, and it pits systems against each other rather than scoring them in isolation. The Q-Score measures factuality minus normalized cost under a competitive match. Across summarization and open-domain QA, single-agent baselines drift into resource-heavy over-optimization, while competition elicits more resource-efficient policies. These gains are small but consistent, and stable across 100 trials. The axis stays discriminative for frontier systems (Gemini-2.5-Pro, and GPT-5 in simulated preview) whose raw factuality scores are already bunched near the ceiling. MAS-HQ provides a reproducible way to measure how much a factual answer costs.

[AI-44] Capacity-Aware Deep Learning for Generalizable Traffic Volume Estimation Across Links and Cities

链接: https://arxiv.org/abs/2607.24056
作者: Léo Hein,Giovanni De Nunzio,Aurélie Pirayre,Laurent Najman
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Network-wide traffic volume estimation typically relies on propagating measurements from fixed sensors, making performance highly dependent on sensor density and limiting deployment in sparsely instrumented networks. We propose a link-level learning framework that estimates hourly traffic volumes from widely available territorial data only, including probe speed profiles, road and topological descriptors, along with weather observations. A supervised local mapping is learned from sparse sensor measurements and evaluated under two generalization settings: intra-network (unseen links within the training network) and inter-network (unseen city). This formulation frames traffic volume estimation as a spatial out-of-distribution generalization problem under sparse supervision. To enhance spatial robustness, we introduce a capacity-aware formulation that models volume as the product of a link-specific structural capacity and an hourly regime-aware utilization ratio, embedding traffic-theoretic constraints directly into the learning process. Extensive experiments in both generalization settings demonstrate that the proposed structural constraints consistently outperform a state-of-the-art baseline under spatial distribution shift.

[AI-45] Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

链接: https://arxiv.org/abs/2607.24054
作者: Jingkun Luo,Da-Tian Peng
类目: Artificial Intelligence (cs.AI)
备注: 17 pages, 3 figures, including supplementary material. Code: this https URL

点击查看摘要

Abstract:A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes intended reasoning from answer acquisition. Outcome evidence and exposure detection do not establish whether success depended on an acquired target; we call this missing evaluation object success provenance. AcquaBench audits it through matched CLEAN, GOLD, and SHAM value substitution on four standardized surfaces with joint qid-clustered analysis. CLEAN retains benchmark-authorized information. GOLD makes the correct target available. SHAM preserves source structure and exposure opportunity but substitutes a matched incorrect value. GOLD minus CLEAN measures the total score response to correct-target availability; GOLD minus SHAM tests whether that response tracks target correctness beyond matched source exposure. In D0, GOLD exceeds SHAM by 19.1 to 25.9 percentage points, showing that success follows the correct value. In D2, GOLD still exceeds SHAM under distributed sufficiency while coloc no longer transfers as a high-score marker, with AUROC 0.376 and 0.142. Behavioral dependence can thus persist beyond this probe’s intended observation unit. In model comparison, a supported 5.0-point CLEAN score gap compresses to a raw GOLD difference of -0.6 points without establishing rank inversion. Agent benchmarks should report success together with whether the evaluated information state supported it.

[AI-46] Quantum-Inspired Evolutionary Neighborhood Search for Arrival-Departure Track Utilization Adjustment under Short-Term Disturbances

链接: https://arxiv.org/abs/2607.24049
作者: Xiaobin Li,Wuming Lei,Yanbin Gao,Weiguang Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Short-term disturbances at major passenger railway stations alter train arrival and departure times as well as the release sequence of station resources. Effective recovery therefore requires coordinated adjustment of arrival-departure track allocation, station resource occupation, and train retiming. This study represents the station resources involved in train arrival, track occupancy, and departure operations as zone-level resource-occupation intervals. An arrival-departure track allocation adjustment model is formulated. Resource compatibility is imposed as the feasibility condition, while train delays and resource reassignment costs are jointly considered. A quantum-inspired evolutionary algorithm combined with neighborhood search (QEA-NS) is proposed to solve the model. Perturbation instances are constructed using GTFS timetable data from Frankfurt Hauptbahnhof, Germany. QEA-NS is compared with CP-SAT under the same candidate resource set and feasibility criteria. Both methods generate solutions satisfying the modeled resource compatibility constraints. QEA-NS yields a total delay of 388 min, compared with 519 min for CP-SAT, representing a reduction of 25.2%. The mean delay of delayed trains decreases from 4.99 to 3.73 min, although QEA-NS requires a longer solution time. Across 10 random perturbation instances, QEA-NS achieves lower total delay in every case. Its mean total delay and standard deviation are 390.5 min and 35.945 min, respectively, compared with 673.8 min and 105.739 min for CP-SAT. The results indicate that, under the adopted resource representation and constraints, QEA-NS improves the delay performance of recovery plans. Its computational efficiency, however, requires further improvement.

[AI-47] he Half-Lives of Generative-AI Evidence: A 40-Record Audit a Claim-Currency Framework and a Reflexive Case of Frontier-Model-Assisted Research

链接: https://arxiv.org/abs/2607.24032
作者: Carlo Iacono(Charles Sturt University, Australia)
类目: Artificial Intelligence (cs.AI)
备注: 20 pages, 6 tables, 1 figure, 1 appendix. Ancillary file this http URL contains the record-level coding. Data cutoff: 17 July 2026; submission-stage frontier note: 27 July 2026

点击查看摘要

Abstract:Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes. First, it audits a maximum-variation purposive corpus of 40 empirical records appearing between 18 July 2025 and 17 July 2026. The audit coded publication route, execution timing, model identity, age of the newest named generation or immutable snapshot, same-family supersession and refresh behaviour. At appearance, the newest named model was a median 281 days old (middle 50%: 75-478; range: 11-939). Median age was 395 days for 25 journal articles, 56 days for 14 preprints and 49 days for one laboratory report. Thirty-five records included a superseded family, seven supplied a precise dated identifier, three clearly refreshed model evidence, and one added a late sensitivity test. All 40 included an OpenAI system, a feature of this corpus rather than a prevalence estimate. The paper distinguishes model age from claim currency and proposes six reporting practices. Second, it treats its own two-day production process as a reflexive case of frontier-model-assisted research creation. GPT-5.6 Sol Pro in ChatGPT supported candidate discovery, source reconciliation, calculations, drafting and critique; the author checked sources, made all substantive decisions and accepts responsibility. This is a proof-of-practice, not a controlled estimate of productivity or quality. By applying its own Model Facts and model-currency statement, the paper shows how rapid AI-assisted research can be made inspectable without treating model output as independent validation. The title uses half-lives metaphorically; no universal decay rate is estimated.

[AI-48] Agent ic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation

链接: https://arxiv.org/abs/2607.24006
作者: Mohan Manivannan,Dalal Alharthi
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注:

点击查看摘要

Abstract:Cloud telemetry arrives at a scale that, paradoxically, makes intrusion understanding harder rather than easier. Attackers operate through legitimate identity, federated session tokens, and cloud native APIs indistinguishable from routine administration, and analysts spend an incident reconstructing context the logs already contain. We present Cloud Decoy AI Agent, a framework pairing a high fidelity cloud decoy with an autonomous language model agent that compresses the path from suspicious activity to an analyst ready report. Connecting a decoy to an agent is not a wiring exercise. The unit of investigation is the session rather than the event, and the session key is obscured by the identity layering federated credentials introduce. The agent’s evidence horizon must be bounded, since an agent free to query full control plane history inherits the cost and false positive profile deception was meant to remove. And cloud telemetry is partly adversary authored, since object keys and user agent strings are attacker chosen values providers record verbatim, which makes any log to prompt path an indirect prompt injection channel that a decoy widens rather than narrows. We address the first two with a session aggregation operator over a pivot tuple drawn only from provider derived fields, and with dynamic prompt generation, a two stage prompt assembly enforcing a grounding invariant by carrying only fields the agent observed. We identify the third as an unaddressed exposure in this class of system, specify the mitigation it requires, and note our prototype does not implement it. Across ten controlled AWS S3 scenarios, nine were reconstructed completely, no report contained an assertion untraceable to an observed artifact, and latency was four to five minutes. We also state what this evaluation does not establish and name the comparisons that would settle it.

[AI-49] Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation

链接: https://arxiv.org/abs/2607.23997
作者: Athanasios G. Papadopoulos
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The ever-growing adoption of Artificial Intelligence (AI) creates the need to deploy Deep Neural Networks in a variety of computational environments. We consider dynamic environments, where computational requirements are subject to change, and we pose the following question: How do we adjust the complexity of an AI classification system, in order to maximize its accuracy, while meeting changing computational constraints? We call this problem Budgeted Image Classification, and we formally formulate it as a resource allocation integer program. Given a computational budget, a batch of images, and a classification system that can make decisions with varying complexity (it has multiple decision points), we explore strategies to allocate images to decision points, in order to maximize accuracy within the available budget. The original integer program is NP-Hard, so, we propose a continuous relaxation, leading to a content-agnostic allocation strategy which assigns images to decision points without considering their particular content. We address this issue by proposing a content-sensitive strategy, that we experimentally show it leads to superior performance. We theoretically study the behavior of our strategies, deriving conditions that must be satisfied by decision points to be suitable for budgeted classification. We analyze fails cases, offering insights for future research directions.

[AI-50] Adaptive Data Admission and Retention for Streaming Federated Learning

链接: https://arxiv.org/abs/2607.23987
作者: Zhuoyi Zhao,Ben Liang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Networking and Internet Architecture (cs.NI)
备注:

点击查看摘要

Abstract:We study streaming federated learning with limited client memory, where newly generated training data incur time-varying sampling costs and must be selectively admitted and retained over time. We consider a joint server-side admission and client-side memory-management framework with the objective of minimizing the cumulative excess population risk under a sampling-cost budget and buffer constraints. We first derive a learning-error bound that explicitly captures the effects of instantaneous training sample size, distinct-sample growth, and reuse imbalance through a characterization of the effective sample size. Through a surrogate penalty obtained from this bound, we develop an Active-Constraint Drift-Plus-Penalty (ACDPP) policy that combines a structured client-side K -step retention rule with a server-side online admission rule and a time-varying rectangular admission region. We further present a sequence of comparison arguments, via an auxiliary constant-admission policy, that connects the ACDPP learning bound to a costless oracle benchmark. This yields explicit guarantees in terms of sublinear regret and sampling-cost violation, while the buffer-occupancy violation is controlled through offline selection of the retention horizon. Experiments on multiple datasets demonstrate that the proposed policy remains close to the oracle benchmark while satisfying the sampling-cost and buffer constraints.

[AI-51] Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

链接: https://arxiv.org/abs/2607.23975
作者: Stefan G. Creadore
类目: Artificial Intelligence (cs.AI); Quantitative Methods (q-bio.QM)
备注: 16 pages, 6 figures, 3 tables. Companion code and data: this https URL . This fork-specific validation study cites, but does not duplicate, arXiv:2510.26887

点击查看摘要

Abstract:Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity. We developed Plato-Bio, a biology-routed extension of the open Plato/Denario architecture that couples explicit workflow states with provenance records, citation checks, claim-to-evidence links, scoped file writes, and publication gates. A source audit identified and repaired three defects that could distort evaluation: loss of task domain in the default factory, omission of declared method signals from scoring, and evidence sidecars that lacked the drafted-claim denominator. On the current clean revision, the full Python suite completed with 931 passes, six skips, and no failures or errors; targeted biology, genomics, evidence/citation, and adversarial-safety suites likewise completed without failure. We evaluated two narrow use cases. In a frozen historical rediscovery task, independent pre-1986 literature bridges ranked the later-studied relation between fish oil and Raynaud phenomenon first; TF-IDF ranked it second and corpus frequency third. This single curated task measures retrospective ranking, not prospective discovery. In a separate comparison of AlphaFold models with experimental structures for 15 human proteins, 11 targets had high-confidence-core C-alpha RMSD below 1 Angstrom (median 0.501 Angstrom). Four targets exceeded 2 Angstrom, and confidence masking reduced the SUMO1 discrepancy from 16.61 to 2.58 Angstrom over 74 residues. The workflow emitted 27 traceable discrepancy regions, all retained as unvalidated hypotheses. Plato-Bio therefore provides reproducible software contracts and auditable screening baselines; broader claims of agent efficacy or biological novelty require preregistered evaluation, independent review, and prospective validation.

[AI-52] Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries

链接: https://arxiv.org/abs/2607.23967
作者: Taeyoung Kim
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 33 pages, 5 figures

点击查看摘要

Abstract:Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight decay, together with a locally quadratic extension to nonlinear neural networks. Our analysis reveals a distinguished population-active component of the empirical null space, which we call the grokking subspace. Along this subspace, the training predictions remain unchanged, leaving weight decay as the sole restoring force and giving rise to a slow dissipative relaxation governed by an exact discrete-time and continuous-time law. We show that only this subspace contributes to the slow asymptotic decay of the population risk and derive explicit iteration-scale predictions for the grokking time, recovering the familiar (1-\beta)/(\eta\lambda) scaling in the weak-regularization regime. The theory further predicts distinct effects of optimizer choice, distinguishing coupled L_2 regularization from decoupled weight decay, and yields causal predictions for interventions that modify the grokking component. We verify all theoretical identities without fitted parameters in a synthetic model where every subspace and relaxation rate is computable in closed form. We further observe genuine delayed generalization in modular addition, where the measured delay follows the predicted scaling and the late-time relaxation agrees closely with the theoretical clock.

[AI-53] EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff

链接: https://arxiv.org/abs/2607.23955
作者: Xiao Ma,Zhiquan Hu,Yi Wei,Chenchen Zhao,Yijun Chen,Jicheng Zhao,Yuming Li Chuang Dai
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups provide no comparative signal and may hide useful search behavior. We present EviBack, an evidence- constrained Teacher backoff that supplies auxiliary super- vision to such groups while preserving verifiable Actor re- wards. It separates evidence assessment from answer refine- ment, preventing reference answers from overriding evidence- insufficiency judgments. A fully automated, end-to-end GPT- 5.5-assisted APE pipeline starts from a manually authored single-prompt dual-task Teacher, automatically partitions and labels rollout data, and performs ablation, task decomposition, evaluation, and selection to produce a gated two-stage Teacher. Compared with the manual design, the resulting Teacher im- proves downstream F1 and valid-answer rate while reduc- ing search, duplicate queries, and forced termination. Across seven open-domain QA benchmarks and three Qwen3 scales, EviBack improves F1 over Search-R1 and raises both single- and multi-hop macro F1. We guarantee that the code will be made publicly available at a later stage.

[AI-54] DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models

链接: https://arxiv.org/abs/2607.23944
作者: Hao Yang,Jin Wang,Xuejie Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusing on question-relevant regions. However, multimodal large language models may deviate from this pattern due to attention drift and the underutilization of visual evidence, which can lead to hallucinations. To mitigate these issues, this study proposes a Dual-Indicator Guided Contrastive Alignment (DICA), which tracks two information-theoretic indicators during inference: Visual Attention Entropy (VAE), which reflects the concentration of visual attention, and Output Image Correlation (OIC), which measures the dependence of generated outputs on the visual input. An abnormal increase in VAE or a decrease in OIC corresponds to different failure modes, which trigger targeted contrastive alignment to restore visual grounding. Experimental results across multiple benchmarks demonstrate that DICA consistently outperforms existing approaches and substantially reduces hallucinations, highlighting the effectiveness of indicator-driven intervention in improving multimodal inference reliability. The code is publicly available at this https URL.

[AI-55] From Cognitive Architectures to Language Agents : A Mechanism-Level Review of Lineage Convergence and Migration Gaps

链接: https://arxiv.org/abs/2607.23942
作者: Haodi Fan,Zucong Lan
类目: Artificial Intelligence (cs.AI)
备注: 28 pages, 9 figures, 14 tables. Review article

点击查看摘要

Abstract:Memory, planning, reflection, and tool use are often compared as feature labels, obscuring the control semantics that determine how an agent actually runs. This review connects ten historical cognitive architectures, eight language-agent runtime families, and forty-two mechanism-focused modern systems. We reconstruct each mechanism through state, control, transition, persistence, failure, learning, and resource governance, then code evidence relation (E1-E4) separately from migration depth (D0-D4). The resulting landscape is uneven. Modern agents have operationalized substantial parts of adaptive memory, failure recovery, dynamic team selection, workflow search, skill induction, resource scheduling, and uncertainty-conditioned action, although often through independent convergence rather than documented inheritance. The strongest remaining opportunities lie in couplings among mechanisms. Closest-baseline screening closes one proposed gap: GraSP already combines calibrated multi-skill selection, typed compilation, verification, bounded repair, and replanning or ReAct fallback. Five residual bundles remain: activation with latency and action utility; typed impasse with isolated substates and resolution compilation; bounded content competition with broadcast and admission learning; persistent intention with reconsideration and live method authority; and uncertainty with resource allocation, interruption, and stopping. We contribute a distinctive-mechanism catalog, an auditable evidence-depth framework, and a falsifiable agenda for testing these bundles as composable runtime invariants.

[AI-56] SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

链接: https://arxiv.org/abs/2607.23933
作者: Yihui Zhang(1),Tianyu Wo(1),Jinghao Wang(1),Xiaoyang Sun(2),Menghao Zhang(1),Cangzhou Yuan(1),Li Li(1),Chunming Hu(1),Albert Y. Zomaya(3),Renyu Yang(1) ((1) Beihang University, (2) University of Leeds, (3) The University of Sydney)
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Performance (cs.PF)
备注:

点击查看摘要

Abstract:As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency. Persistent long-lived sandbox reservations incur excessive memory overhead at scale, while lazy on-demand instantiation generates severe cold-start penalties that degrade response performance under multi-tenant, multi-turn agent workloads. To resolve this dilemma, we present SpecBox, a runtime built around speculative sandbox preallocation tailored for dynamic LLM agent execution pipelines. At its core, SpecBox implements keyword matching and streaming semantic embedding to enable intent-driven sandbox prewarming, which identifies pending tool execution demands mid-LLM token generation and fully overlaps sandbox bootstrapping with model inference. To extend prewarming windows across sequential agent steps, the framework leverages context-aware stochastic prefetching atop a sandbox dependency graph to probabilistically forecast future sandbox switches ahead of execution. We complement these speculative mechanisms with two orthogonal optimizations: a semantic result cache that prunes redundant repeated sandbox invocations, and a dedicated out-of-band shared-memory transport plane that bypasses conventional network serialization to deliver zero-copy artifact transfers. Evaluated on high-concurrency multi-turn agent traces, our prototype demonstrates that SpecBox cuts P99 end-to-end latency by up to 2.9\times relative to the on-demand sandbox baseline, while slashing peak memory consumption by 45.9% compared to permanently reserved sandbox deployments. Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Performance (cs.PF) Cite as: arXiv:2607.23933 [cs.DC] (or arXiv:2607.23933v1 [cs.DC] for this version) https://doi.org/10.48550/arXiv.2607.23933 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-57] MemTX: Transactional Belief Commit for Stateful Agent Memory

链接: https://arxiv.org/abs/2607.23929
作者: Xiaoyang Li,Yiqi Wang,Haohui Lu,Zhi Chen,Mo Li,Pingan Song,Taotao Cai
类目: Artificial Intelligence (cs.AI)
备注: Preprint

点击查看摘要

Abstract:LLM agents increasingly coordinate through persistent shared memory: one agent’s write becomes another agent’s premise, and eventually a tool call with real side effects. Current agent memory systems treat every accepted write as immediately actionable truth, so a polluted tool result, a stale update, or a teammate’s half-finished note can silently drive an irreversible action. We argue that a memory write is not a belief commit. We present MemTX, a transactional belief-commit protocol. Each record carries evidence, permissions, provenance, and validity. Writes are staged inside snapshot-isolated transactions and admitted by a validate-and-commit pipeline, irreversible tool calls are gated on in-flight belief state, and retracting a belief triggers typed cascading repair of its derived records and tool side effects. Two invariants, action-safety gating and cascade-repair completeness, are machine-checked by property-based testing and bounded exhaustive enumeration of 5.5 million protocol states, with zero violations. Across five backbones from three model families, MemTX leads all eight baselines with paired-McNemar significance on four backbones and statistically ties the best baseline on the fifth and strongest, while remaining the only method with zero downstream harm on every backbone. Backbone capability does not substitute for commit discipline.

[AI-58] GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models

链接: https://arxiv.org/abs/2607.23913
作者: Jun Ling,Tao Huang,Junzhuo Liu,Bowen Tang,Peng Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream language-model inference cost. Existing token-reduction methods assess token utility through token-wise importance, query relevance, coverage, pairwise diversity, or subset-level objectives. Our key insight is to view visual token reduction through selected-span complementarity: instead of scoring a token in isolation or through pairwise relations, we assess how much of its feature is orthogonal to the span of the already retained subset. Based on this perspective, we propose Greedy Orthogonal Token Selection (GOTS), a training-free and query-agnostic method. At each step, GOTS selects the token with the largest residual energy orthogonal to the current retained span. This rule exactly maximizes the one-step augmented Gram determinant among candidate additions, giving each greedy step a precise local geometric guarantee for subset expansion. Across five high-resolution VLM backbones from the Qwen-VL and InternVL families and eleven diverse benchmarks, GOTS achieves higher average performance retention than the strongest evaluated baselines, and a controlled OCRBench study shows that it reduces model-side time-to-first-token after accounting for selection overhead. Code is available at this https URL.

[AI-59] Embodied GPT -5.1: Evidence of a World Model?

链接: https://arxiv.org/abs/2607.23899
作者: Roberto Spinelli,Thiago C. Martins
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 6 pages, 16 figures. Published in the 2026 Brazilian Conference on Robotics (CROS)

点击查看摘要

Abstract:This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mobile robot despite having no prior embodiment, no training in simulated environments, and no exposure to sensorimotor experience. Using only low-resolution first-person images and a discrete action set, the model was tasked with navigation and object-directed behaviors such as locating and contacting a target toy. Across multiple trials, GPT-5.1 demonstrated emergent capabilities that suggest elements of spatial reasoning and physical understanding. These included maintaining short-term memory of object locations after they left the camera frame, inferring the physical consequences of its own movements, and executing coherent action sequences such as colliding with an object and reversing to visually verify the outcome. At the same time, the model displayed inefficiencies and perceptual limitations, including imprecise alignment strategies and occasional misidentification of distant distractors. Overall, the results indicate that GPT-5.1 exhibits signs of world-model-like behavior in an embodied setting, despite the absence of any embodiment-related training, a finding that challenges long-standing views in cognitive science and robotics which hold that a physical body is a necessary prerequisite for developing such forms of intelligence. The findings motivate deeper investigation into the emergence, limits, and robustness of physical understanding in large language models.

[AI-60] Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery

链接: https://arxiv.org/abs/2607.23896
作者: Debajyoti Ray,Niranjan Srinivas
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathway merits optimization. We formulate this as a sequential decision problem with a discrete pathway-identification stage and a continuous within-pathway optimization stage under heterogeneous experimental costs. Our implementation, Coactive learning, combines a cost-sensitive Bayesian hypothesis-discrimination policy motivated by EC2 (Golovin et al., 2010) with Gaussian-process Bayesian optimization (Srinivas et al., 2010). Under explicitly stated assumptions, the expected spend of one fixed-budget campaign attempt is bounded by the expected pathway-identification cost plus the capped within-pathway optimization budget. We evaluate the method on synthetic benchmarks constrained by selected results reported for PNNL’s CICERO selective-precipitation study (Ritchhart et al., 2026). The method performs comparably to an oracle-pathway Bayesian-optimization reference and to a strong split-plate baseline that discriminates pathways with its first plate, without receiving an oracle label for the correct pathway. It is given a candidate hypothesis space and a diagnostic likelihood model. On an NdFeB-inspired instance, it avoids the simulated penalty of a commit-first baseline that initially selects a plausible but inferior hydroxide pathway. This hypothetical wrong-first-commitment scenario is motivated by the hydroxide-oxalate performance contrast reported by CICERO. We characterize the sensitivity of these conclusions to the assumed cost model. The code and benchmark are open source.

[AI-61] Physics-Informed Neural Networks for Predicting Nitrous Oxide Flux

链接: https://arxiv.org/abs/2607.23880
作者: Freddy Yu,Jashanjeet Kaur Dhaliwal,Subhadeep Chakraborty
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注: 28 pages, 7 figures

点击查看摘要

Abstract:Nitrous oxide (N _2 O) is the dominant ozone-depleting substance emitted in the 21st century, and the third largest contributor to anthropogenic greenhouse gases due to its high potency and long atmospheric lifetime, with more than 70% of N _2 O emissions occurring as a result of agricultural processes. Current approaches to predicting N _2 O flux emissions include process-based models such as DayCent and Cycles, as well as classical AI models, but the application of Physics-Informed Neural Networks (PINNs) to predicting N _2 O flux emissions is largely underexplored. Our paper draws upon the mechanistic equations that underlie the DayCent family of process-based models to construct a rigorously derived, literature-traceable physics residual. We then build and train an MLP-based PINN on a multi-site agricultural dataset spanning four geographically distinct US agricultural sites. Across all tested values of the physics loss weighting hyperparameter \lambda , our PINN consistently and substantially outperformed uncalibrated Cycles simulation (R ^2=0.01 ), with our MLP baseline achieving mean R ^2=0.411 across ten random seeds. Physics constraints consistently degrade model performance in holdout validation, with marginal degradation at low \lambda and significant degradation at high \lambda , but consistently improve model performance and reduce performance variability in leave-one-site-out validation. This suggests that physics constraints sacrifice in-distribution accuracy for out-of-distribution robustness, anchoring the model toward biogeochemically plausible behavior on unfamiliar soil conditions — though cross-site generalization remains challenging, with negative R ^2 across all seeds and \lambda values on our geographically distinct held-out site.

[AI-62] A Coulomb Particle Model for Learning Kernel Attention in Transformers ICML2026

链接: https://arxiv.org/abs/2607.23869
作者: Masoud Badiei Khuzani,Sharath Honnaiah,Atiq Islam,Alex Cozzi,Abraham Bagherjeiran
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Workshop on High-dimensional Learning Dynamics (HiLD), ICML 2026

点击查看摘要

Abstract:Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution. We propose a particle-based method that learns this distribution by optimizing kernel-target alignment while regularizing particles with a Riesz/Coulomb repulsive potential. The resulting Hamiltonian yields diverse, task-adaptive random features and admits a mean-field description through a McKean–Vlasov equation. We instantiate the method in linearized Transformer attention by learning positive random-feature maps in a first alignment phase, then freezing the kernel and training the remaining network parameters with cross-entropy. Experiments on synthetic classification and sentence-level benchmarks show that learned kernelized attention can improve accuracy, calibration, and robustness for several feature maps while preserving linear-attention inference complexity.

[AI-63] Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search

链接: https://arxiv.org/abs/2607.23854
作者: Haijiang Yan,Jian-Qiao Zhu,Liqiang Huang,Ming Meng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms. In the Euclidean traveling salesman problems (TSP), people rapidly produce tours that are near-optimal, despite severe limits on time and computation. What makes a tour human-like, and how might such solutions be learned? Here we address these questions through a large-scale behavioral and computational investigation of human performance in Euclidean TSP. We sampled a broad space of TSP instances, collected human solutions, and compared them with neural policies based on Pointer Networks, which are recurrent neural networks with an attention-based pointing mechanism that define probability distributions over valid tours. We trained these networks under multiple objectives, including reinforcement learning (RL), supervised learning from optimal tours, supervised learning from human tours, and RL fine-tuning after optimal-supervised pretraining. Human tours were not identical to optimal tours, but occupied a near-optimal geometric basin: they shared many structural properties with optimal solutions while preserving systematic human-specific deviations. The best account of human tours was not direct imitation of optimal tours, but a model pretrained on optimal tours, fine-tuned by RL, and decoded through \textBest-of-N sampling. These findings suggest that human-like solutions may emerge from a combination of structured supervised learning, RL, and test-time search, echoing computational principles underlying many modern artificial intelligence systems.

[AI-64] Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach

链接: https://arxiv.org/abs/2607.23845
作者: Anh Ngo,Nicolas Rollet,Catherine Pelachaud,Chloé Clavel
类目: Artificial Intelligence (cs.AI)
备注: Accepted to ICMI 2026 (International Conference on Multimodal Interaction), October 5-9, 2026, Napoli, Italy

点击查看摘要

Abstract:Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a recipient signals a problem in speaking, hearing, or understanding, prompting the previous speaker to resolve it. In the case of conversational agents, it is essential to accurately identify these repair initiation strategies to address communication breakdowns efficiently. While conversational analysis studies have shown that OIR initiation is accompanied by both verbal and non-verbal signals such as gaze shifts, facial expressions, body postures, and hand gestures, existing computational approaches rely mainly on text and audio. This paper introduces a novel multimodal model for OIR detection and classification, incorporating a set of visual features drawn from conversation analysis. We evaluate our approach on two corpora with distinct languages and interaction settings. Results demonstrate that visual information consistently improves performance over text and audio baselines, and provide insights into cross-modal feature contributions across two corpora.

[AI-65] Limbomorphs

链接: https://arxiv.org/abs/2607.23842
作者: Alex Alvarez,Michael Levin
类目: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI)
备注: 3 pages, 3 figures. Extended abstract accepted as a Late Breaking Abstract at the Artificial Life Conference (ALIFE 2026)

点击查看摘要

Abstract:Artificial life systems are typically defined by a set of dynamical rules over an environment, an agent, or both, from which lifelike patterns may emerge. Gifbreeder is an animated version of the interactive evolutionary computation (IEC) platform Picbreeder, and was initially created to generate visual art. Instead of encoding the agent or the environment, Gifbreeder genomes encode a spatiotemporal field and evolve through the user’s aesthetic selection. The evolved expressions can sometimes resemble motile lifelike creatures that we term Limbomorphs, given that they exist in a deterministic three-second looping “limbo”. We assess their behavior via input-space perturbations and find species-specific reactions to different kinds of perturbations. We discuss whether these reactions may reflect goal-directed behavior like navigation, or merely the appearance of it, and more broadly how agent-like dynamics may emerge in a system with no explicitly defined agent, environment, or interaction rules.

[AI-66] ACM: Agent ic Context Management for Long Horizon Tasks

链接: https://arxiv.org/abs/2607.23809
作者: Xiaochuan Li,Ryan Ming,Meng Chu,Shuai Shao,Rong Jin,Chenyan Xiong
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existing context compression methods inevitably incur information loss and are triggered by rigid heuristic rules, leaving them misaligned with the agent’s evolving reasoning focus. We propose Agentic Context Management (ACM), a framework that equips agents with purpose-built context editing tools for lossless context management. Inspired by the interaction between short-term and long-term human memory, the agent autonomously decides when to compress its context, offloads discarded content to an external memory system, and queries it on demand for later retrieval. Building on this framework, we further develop a post-training pipeline that constructs high-quality demonstrations of context management and improves model performance on both agentic search and coding tasks. Further analysis reveals that effective context management reduces peak token pressure, enables extended explorations, and yields more consistent solutions across independent trials. Code, data, and model checkpoints are available at this https URL.

[AI-67] From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

链接: https://arxiv.org/abs/2607.23802
作者: Qinsi Wang,Jing Shi,Huazheng Wang,Kun Wan,Yiran Wu,Bo Liu,Qingyun Wu,Hai Helen Li,Yiran Chen,Handong Zhao,Wentian Zhao
类目: Artificial Intelligence (cs.AI)
备注: COLM 2026

点击查看摘要

Abstract:Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verified. Open-ended tasks instead often rely on human preferences, reward models, or LLM-based judges, introducing evaluation bias, judge capability bottlenecks, and additional inference this http URL on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals. We instantiate RLSVR with SpyRL, a multi-agent self-play environment inspired by Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality. Experiments on text summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods on non-verifiable tasks and yields consistent gains on verifiable reasoning tasks. These results demonstrate that task transformation can extend scalable RLVR-based self-improvement beyond inherently verifiable domains. Models and code have been released at this https URL.

[AI-68] Maximum Satisfiability of Simple Temporal Problems

链接: https://arxiv.org/abs/2607.23785
作者: Johannes K. Fichte,Johanna Groven,Peter Jonsson,Victor Lagerkvist,Jorke M. de Vlas
类目: Computational Complexity (cs.CC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The Simple Temporal Problem (STP) is a core framework for quantitative temporal constraints. As STP data can be inconsistent, we study MAXSTP: compute a maximum-cardinality consistent subset of constraints. This extension is NP-hard, and we analyze its parameterized complexity under measures that capture practically relevant instance features: the number of variables n (instance scale), the maximum coefficient magnitude k (numeric range), and structural parameters of the constraint graph such as treewidth tw (decomposability) and vertex cover size vc (density). We show that MAXSTP is W[1]-hard parameterized by n , implying that n and parameters that depend on n (including tw and vc ) are insufficient for fixed-parameter tractability. For combined parameters, we give an O^(k^n) -time algorithm, yielding single-exponential solvability for fixed k . While k+tw remains W[1]-hard, MAXSTP is in XP via an O^((n\cdot k)^tw) algorithm. Our results suggest that MAXSTP is often computationally harder than optimizing qualitative CSPs. We verify that many such problems (including RCC-8 and Allen’s algebra) are FPT when parameterized by n or tw . However, we also demonstrate that FPT algorithms for MAXSTP are indeed possible but with other parameters such as k + vc .

[AI-69] A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

链接: https://arxiv.org/abs/2607.23784
作者: Daphne Chen,Archit Ritesh Jain,Eric Goossen,Emma Romig,Michael Murray,Nick Walker,Maya Cakmak
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail. In this work, we propose ARCHITECT, a framework that treats robot policy acquisition as an interactive program synthesis task. ARCHITECT leverages the reasoning capabilities of LLM coding agents to synthesize modular robot programs that utilize a suite of perception and control tools. Unlike end-to-end models where distribution shift leads to unpredictable, cascading failures, our modular architecture allows users to isolate failures and localize feedback at the level of abstraction required. We introduce an iterative process where a human supervisor provides natural language corrections to steer the policy. These corrections are grounded in the policy code by program execution traces and distilled into a persistent skill library, a form of long-term in-context learning which enables the agent to accumulate a repertoire of reusable, interpretable behaviors. In a benchmark evaluation on a Franka Panda robot, ARCHITECT outperforms state-of-the-art VLA models and program synthesis baselines on complex, long-horizon tasks, including articulated object manipulation and cloth folding. Our results demonstrate that the synthesized skill library enables the system to transfer to novel tasks with decreasing human intervention, providing a steerable and data-efficient alternative to black-box robot learning. Website: this https URL

[AI-70] Scale Weight Decay and Train Better

链接: https://arxiv.org/abs/2607.23777
作者: Anuj Apte
类目: Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注: Comments welcome

点击查看摘要

Abstract:The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a constant decoupled weight decay which causes the network weights to shrink steadily over the course of training. Taking inspiration from the Robbins–Monro conditions, we propose to scale weight decay by the fraction of the peak learning rate \eta/\eta_\max . We prove that this scaled weight decay preserves the asymptotic stationarity guarantees of the corresponding unregularized methods for both stochastic gradient descent and the non-Euclidean spectral optimizer Muon, thereby avoiding the additional asymptotic bias introduced by constant decoupled weight decay. This retains the stability benefits of weight decay without changing the asymptotic optimization target. Using a steady-state analysis, we explain why under standard weight decay the weight norm shrinks steadily as training proceeds, whereas under scaled weight decay it settles to a roughly constant value. When applied to the training of mixture-of-experts models, Muon with scaled weight decay (Muon-SW) consistently outpaces Muon with identical hyperparameters, reaching the same validation loss \mathbf30% faster at our largest scale across models from 72 - 930 million parameters trained at \sim 600 tokens per active parameter. If this trend continues to hold, the method promises to substantially accelerate the pre-training of frontier models while requiring only a few lines of code to implement.

[AI-71] Outcome-Fair Restless Multi-Armed Bandits for Stochastic Deadline Scheduling

链接: https://arxiv.org/abs/2607.23772
作者: Shakti Sharma,Rahul Meshram
类目: ystems and Control (eess.SY); Artificial Intelligence (cs.AI)
备注: 8 pages, 4 figures, conference

点击查看摘要

Abstract:We study a restless multi-armed bandit (RMAB) problem for a stochastic deadline scheduling application. RMAB problems are solved using the Whittle index policy. The goal in RMAB is to maximize the expected cumulative discounted reward maximization. The Whittle index policy maximizes reward, but is not fair among two classes. In this paper, we introduce fairness criteria and study an outcome-fair model for RMAB which allows fairness for jobs and users structurally disadvantaged demographic classes. We formulate an outcome fair stochastic deadline scheduling problem as RMAB, and we develop the outcome fair Whittle index policy. We define a virtual queue mechanism that dynamically enforces long-term completion rate guaranties across demographic groups. We analyze a standard Whittle index policy and the outcome-fair index policy. We demonstrate the performance of our algorithms with numerical examples. We compare policies—Whittle index policy (no fairness), input-fairness Whittle index policy, outcome fair Whittle index policy. We observe that the outcome-fair Whittle index policy provides better fairness among classes compared to other policies. We demonstrate a trade off between fairness and profit. This decreases as the server capacity increases. Comments: 8 pages, 4 figures, conference Subjects: Systems and Control (eess.SY); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.23772 [eess.SY] (or arXiv:2607.23772v1 [eess.SY] for this version) https://doi.org/10.48550/arXiv.2607.23772 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-72] raining Language Models to Cooperate with Inference-Time Controllers

链接: https://arxiv.org/abs/2607.23771
作者: Moumita Choudhury,Vanshaj Khattar,Jing Liu,Toshiaki Koike-Akino,Ankush Chakrabarty,Shlomo Zilberstein,Ye Wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language model (LLM) performance increasingly depends not only on the base model, but also on the inference-time controller used to organize reasoning. Existing post-training methods, however, typically optimize for a single fixed interaction pattern, despite real deployments relying on diverse controllers such as Chain-of-Thought, self-consistency, debate, planning, and verification pipelines. This creates a training–deployment mismatch and limits transfer to new workflows. We introduce CALM (Controller-Aware Language Models), a post-training framework that explicitly places controllers in the training loop. We formulate controller-aware post-training as multi-task reinforcement learning over controller-induced interaction protocols, where controllers are compositions of reusable local reasoning modules. This structure also induces a module-level decomposition of mixed-controller training under a turn-level GRPO objective, enabling a systematic study of controller and module-aware training strategies. We evaluate CALM on held-out controller compositions and broader controller shifts, showing that controller-aware post-training improves generalization across inference-time workflows beyond single-controller optimization.

[AI-73] WISERouter: LLM Routing with Workload Budget Constraint

链接: https://arxiv.org/abs/2607.23765
作者: Yifei Li,Zihui Gao,Laks V.S. Lakshmanan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale. LLM routing exploits diversity in model capability and cost by assigning each query to a suitable model to balance utility and budget. Current methods have two limitations: (i) they either use heuristics that do not always enforce the budget constraint or impose a fixed per-query budget that cannot adapt across the workload and leads to suboptimal performance; (ii) they require supervised learning on a dense dataset with statistics for every query-model pair, which is expensive to collect. To address these challenges, we formulate LLM routing as a constrained contextual multi-armed bandit problem and introduce WISERouter (WR for short), a framework that supports offline learning from historical interactions as well as online learning with exploration. We further prove that WR-Online achieves a sublinear regret bound of O(\sqrtT) over a time horizon T . Empirical results on RouterBench and SWE-Bench demonstrate that (i) WR-Offline surpasses existing baselines in performance under a fixed budget and adheres more closely to budget constraints, and (ii) WR-Online achieves comparable performance to the baselines, while using substantially less exploration data.

[AI-74] AI Strategy: How to Choose What AI Product to Implement

链接: https://arxiv.org/abs/2607.23733
作者: Foster Provost,Panos Ipeirotis
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); General Economics (econ.GN); Applications (stat.AP)
备注: 27 pages, 2 figures, 1 table. Submitted to Big Data (SAGE)

点击查看摘要

Abstract:Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve opposite decisions. At the residential real-estate brokerage Compass, one AI product (Likely-to-Sell recommendations) flagged sales outreach opportunities and went on to account for nine figures in annual gross commission revenue. Another championed AI product (a Time-on-Market pricing tool) was rightly shelved. A simple ROI estimate could not distinguish the two. We present expected ROI (eROI), a framework that decomposes each bet into three components and rates them separately: Value if Successful, Likelihood of Success, and Investment Required. Each maps to a question executives can answer before building: How valuable would it be if it worked? How likely is it to work? And what would it cost to implement? Separating the three breaks a common catch-22: teams cannot estimate ROI until they know whether a project will work, yet cannot know whether it will work without building it. Judging Value if Successful on its own dissolves the loop, letting a team argue that a product would be valuable if it worked while it weighs how likely that is. The framework also asks, before ranking anything, whether there are enough good ideas on the table. After ranking, it guides assembling a portfolio of bets rather than funding only the single top-ranked project. We illustrate eROI on Compass’s candidate AI products. Precise ROI estimates are hard to make given the inherent uncertainty of AI projects. Coarse business-level ratings of the three components are enough to tell strong bets from weak ones.

[AI-75] E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

链接: https://arxiv.org/abs/2607.23722
作者: Weihuang Zheng,Tianyuan Zou,Eileen Ye,Alphet Liu,Youyong Kong,Ya-Qin Zhang,Duran Zheng,Maxm Pan
类目: Artificial Intelligence (cs.AI)
备注: 29 pages, 14 figures, 6 tables

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes. We refer to this capability as multi-step tool use. Existing benchmarks have advanced tool-use agent evaluation, but often focus on isolated API calls, short trajectories, or settings that are difficult to scale or control. We introduce E-Bench, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting. E-Bench decouples environment synthesis from task synthesis: graph-guided database filling builds reusable, orphan-free product environments, while generator-solver asymmetry creates tasks with both an information gap and a tool gap, requiring agents to discover hidden data and compose multiple tool calls before changing state. Outcomes are graded deterministically by database-state diffs. Since both environments and tasks are synthetic, E-Bench is controllable at the environment level and scalable at the task level. Benchmarking 11 cutting-edge LLMs shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code extension, reliability (Pass^3) remains below 70%.

[AI-76] Offline-Online Curriculum RL for Multimodal Reasoning

链接: https://arxiv.org/abs/2607.23700
作者: Wendi Deng,Hang Du,Guoshun Nan,Haokun Tian,Jiaqi Yu,Xinlei Cao,Jaile Li,Jingfeng Chen,Ling Deng,Ting Li,Hao Yang,Jun Liu,Xudong Jiang,Sicong Leng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines interpretability and reliability, suggesting reliance on spurious shortcuts rather than faithful reasoning. Although efforts have explored step-level supervision, distinguishing decisive steps from redundant ones remains challenging. We propose O^2 -CritiCuRL, a novel curriculum reinforcement learning framework that introduces critical-step awareness through an iterative offline-online paradigm. In the offline stage, O^2 -CritiCuRL conducts multi-rollout analysis over step-annotated trajectories to estimate step-level importance, allowing the framework to distill critical reasoning steps and filter out redundant ones. In the online stage, we employ a progressive step-level reinforcement learning strategy, where truncated chains guide the model to infer missing steps and refine its reasoning, thereby sharpening its focus on critical steps and overcoming the limitations of static supervision. Extensive experiments on multimodal reasoning benchmarks show that our method achieves state-of-the-art performance while delivering superior training and inference efficiency. Code is available at this https URL.

[AI-77] Offline-to-Online Creative Optimization with Generative Models and Adaptive Testing

链接: https://arxiv.org/abs/2607.23696
作者: Kevin Lee,Benjamin Letham,Zhiyuan Jerry Lin,Elodie Samson,Eric Onofrey,Poppy Zhang,Shawndra Hill,Eytan Bakshy
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Ad creative optimization is increasingly constrained by evaluation rather than generation. Generative models can produce many plausible creatives, but reliable evaluation requires online experiments, in which only a limited slate can be tested. We study how to use data from historical A/B tests to generate and select the candidates in that slate. We developed and deployed a performance-driven offline-to-online workflow that guides creative generation with a predictive model as an inference-time critic. In the offline phase, we use a predictive model trained on historical experiments to rank and refine variants created by a generative model. A final test slate is then deployed in an online adaptive experiment. In a 50-arm field experiment, we found that the best creative generated with this method yielded 45.1% higher engagement than the best human-authored creative. Two additional experiments showed the same upper-tail pattern, with lifts of 46.7% and 36.2%. We found that despite the predictive model being too noisy to directly identify the best creative offline, it effectively guides the generative model toward creating strong candidates that can be efficiently evaluated in an adaptive experiment. The results suggest a design principle for creative optimization with generative models: use predictive models to guide generation of a slate to test, judge the slate by whether it contains high-performing candidates at a feasible test size, and use adaptive experiments to select among candidates while limiting traffic lost to weak arms.

[AI-78] Compute Globally Materialize Locally: The Memory Contract of Sparse Event-KV

链接: https://arxiv.org/abs/2607.23693
作者: Zefeng Cai,Zerui Cai
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Eviction and episodic-memory schemes therefore rest on a premise rarely tested directly, that a retained event is still informative once the observations that produced it are gone. We test it by omitting one earlier observation from what is served, across otherwise identical agent histories. Among items sensitive to that observation, the answer overwhelmingly follows the omitted value, though no served span says which value is correct. We call this semantic materialization: a downstream event’s cached rows act as an independently servable view of computation whose inputs are gone. It can also be written on purpose. A deliberately phrased, answer-free event raises donor-aligned recovery from 6% to 51% on Qwen3-8B without ever naming the value, whereas passively harvesting natural mentions from long-term dialog yields no detected advantage. What such a row carries is specific and bounded. Compact state survives, larger payloads decay toward chance, and whether a construction writes at all turns on phrasing rather than on meaning alone, so two phrasings the model comprehends equally well can diverge sharply. The result is a memory contract for sparse event-KV serving: what to write, where it lands, and what survives once the source is gone. For anyone who evicts the corollary is that dropping a source event and observing no accuracy loss does not show the source was unnecessary.

[AI-79] Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems

链接: https://arxiv.org/abs/2607.23678
作者: Mingzhou Fan,Siyuan Xu,Mingxuan Yuan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) enable autonomous agents for reasoning, planning, and tool use. Recent systems increasingly organize these agents as graphs of specialized, interconnected nodes. Although graph-based orchestration supports flexible decomposition and coordination, it creates a key challenge: \textbfattention allocation. As workflows grow, existing approaches often execute graph components uniformly, wasting resources on irrelevant or low-impact tasks. We introduce \textbfAttention Orchestration, a paradigm that extends Transformer-style attention from token representations to workflow-level agent coordination. Our framework, \textbfAdaptive Goal-aware Attention Orchestration (AGAO), dynamically estimates agent importance based on user objectives, graph dependencies, and computational constraints. AGAO combines three components: (1) goal-aware attention, measuring semantic relevance between user goals and agent capabilities; (2) topology-aware attention, modeling structural dependencies in agent graphs; and (3) resource-aware attention, allocating budgets and execution priorities across heterogeneous agents. Together, these mechanisms transform static agent graphs into adaptive systems that focus computation on goal-critical reasoning paths. Experiments across diverse multi-agent workloads show that AGAO improves task effectiveness while reducing unnecessary computation, latency, and token consumption compared with existing graph-based execution strategies. Our work establishes \textbfAttention Engineering as a direction for scalable, intelligent multi-agent systems. Code: this https URL.

[AI-80] SpecAHD: Localize to Specialize for Automated Heuristic Design in Large-Scale Routing Problems

链接: https://arxiv.org/abs/2607.23676
作者: Kezhao Lai,Yutao Lai,Hai-Lin Liu
类目: Artificial Intelligence (cs.AI)
备注: 9 pages, 1 figure, and 3 tables. Supplementary material included

点击查看摘要

Abstract:LLM-based automated heuristic design (AHD) typically scores executable programs on complete instances or within fixed solver components. In large-scale routing problems, localized reconstruction reduces the size of each optimization task, but repair regions within the same incumbent can exhibit substantially different structures. One construction rule must therefore compromise across them. In this paper, we propose SpecAHD, a coupled bilevel framework for within-instance specialization. An upper-level search learns where to expose bounded repair regions, while a lower-level search evolves a complementary repertoire of executable heuristics for the induced repair tasks. The upper-level program determines the repair tasks seen by the lower level, while checked repair outcomes determine how upper-level programs are evaluated. The lower-level objective favors heuristics that perform well on average or solve tasks that the current repertoire handles poorly. For the repair tasks induced by a fixed upper-level program and a fixed lower-level candidate pool, this objective is monotone submodular, allowing greedy repertoire selection with a (1-1/e) approximation guarantee. Across four routing problems and multiple LLM backbones, SpecAHD reduces held-out objective cost by up to 57.7% against the strongest competing AHD baseline and outperforms the per-instance baseline envelope on most public instances.

[AI-81] CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation

链接: https://arxiv.org/abs/2607.23647
作者: Gengyu Zhan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse enduring preferences, transient intent, and exposure-induced behavior into one profile. This makes recommendation vulnerable to feedback loops: repeated exposure is mistaken for preference, immediate clicks dominate delayed satisfaction, and fluent explanations need not reflect the ranking decision. We propose our method, a model-agnostic framework for long-horizon recommendation. Our method uses a frozen multimodal language model to convert item content and feedback into evidence-grounded semantic atoms, then maintains separate short-term, long-term, and exposure memories. Propensity-weighted updates reduce policy-induced exposure bias, while a conservative offline critic reranks candidates for delayed satisfaction under a behavior-support constraint. Explanations use only influential evidence atoms and are checked by counterfactual deletion. We provide an identification result and evaluate the framework in e-commerce-like, news-like, and short-video-like environments. Across ten seeds, our method improves discounted long-term value over the strongest alternative by 6.1%, 7.6%, and 6.7%, respectively. Twenty-seed paired ablations show significant value drops after removing propensity correction (0.739 +/- 0.191) or conservative support regularization (0.523 +/- 0.234). A frozen instruction language model also more than doubles semantic-atom NDCG over TF-IDF on a held-out paraphrase benchmark.

[AI-82] Extending Desbordante with Probabilistic Functional Dependency Discovery Support

链接: https://arxiv.org/abs/2607.23636
作者: Ilia Barutkin,Maxim Fofanov,Sergey Belokonny,Vladislav Makeev,George Chernishev
类目: Databases (cs.DB); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Data profiling aims to extract complex patterns from data for further analysis and use that data in domains such as data cleaning, data deduplication, anomaly detection, and many more. Functional dependencies (FDs) are one of the most well-known patterns. However, they are poorly suited for these tasks, as real data is usually dirty, and the rigid definition of FDs does not allow algorithms to locate them. For this reason, there are several formulations aimed at relaxing FDs to support dirty data, with approximate functional dependency (AFD) being the most popular one. Another formulation is the Probabilistic Functional Dependency (pFD), which we aim to support inside Desbordante - a science-intensive, high-performance and open-source data profiling tool implemented in C++. However, pFDs are relatively poorly studied, compared to AFDs. In this paper we study pFDs, both analytically and empirically. We start by assessing how different pFDs and AFDs are by studying cases in which pFDs have an edge over AFDs. Then, we implement the algorithm for pFD discovery, as well as study its run time and memory consumption. We also compare it with an AFD discovery algorithm. Lastly, we study the output of both algorithms to learn whether or not it is possible to use AFD discovery algorithm to get pFDs and vice versa. Subjects: Databases (cs.DB); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG) ACMclasses: H.3; I.5; J.0 Cite as: arXiv:2607.23636 [cs.DB] (or arXiv:2607.23636v1 [cs.DB] for this version) https://doi.org/10.48550/arXiv.2607.23636 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: 35th Conference of Open Innovations Association (FRUCT), Tampere, Finland, 2024, pp. 158-169 Related DOI: https://doi.org/10.23919/FRUCT61870.2024.10516409 Focus to learn more DOI(s) linking to related resources

[AI-83] Variational-Ising-Attention (VIA):TailoredAttentionMattersfor Science

链接: https://arxiv.org/abs/2607.23634
作者: Rui Wang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Chemical Physics (physics.chem-ph)
备注: 13 pages, ~20 figures

点击查看摘要

Abstract:Attention enables context modeling via query-key scoring with softmax normalization. Driven by industrial long-context demands, mainstream research has converged toward sparsity and efficiency–yet softmax’s independence assumption persists. For scientific tasks unburdened by long-token constraints, however, richer structured coupling may often be essential, making tailored attention both viable and more appropriate. To this end, we propose Variational-Ising-Attention (VIA), which augments softmax normalization with an interacting Ising model; attention patterns emerge from learnable pairwise couplings via variational mean-field inference, redefining attention from a ranking over isolated items to a collective state over interacting entities. We instantiate VIA on retrosynthesis reaction center prediction, a task inherently governed by cooperative bond-breaking constraints. Comprehensive experiments across model variants, coupled with mechanistic analyses, demonstrate that VIA consistently and substantially outperforms standard softmax attention. More broadly, our findings suggest that for scientific problems, the optimal solution is not general-purpose efficiency, but appropriately tailored attention aligned with intrinsic domain structure. This work provides a theoretically grounded and empirically validated instantiation of this paradigm.

[AI-84] Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms

链接: https://arxiv.org/abs/2607.23632
作者: Yakov Kuzin,Dmitriy Shcheka,Michael Polyntsov,Kirill Stupakov,Mikhail Firsov,George Chernishev
类目: Databases (cs.DB); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Performance (cs.PF)
备注:

点击查看摘要

Abstract:Science-intensive data profiling focuses on discovery and validation of various patterns in datasets. This study considers discovery of one such pattern - order dependency (OD). Simply put, OD states that some list of columns is ordered according to another one. It is of use for database query optimization, data cleaning and deduplication, anomaly detection, and much more. Existing discovery methods have approached this problem solely from the algorithmic standpoint, without focusing on the implementation side. At the same time, this problem is very computationally intensive, and therefore this part should not be ignored, as it brings ODs closer to industrial use. In this paper, we study two algorithms for OD discovery which target different OD axiomatizations - FASTOD and ORDER. We start by reimplementing these algorithms in C++ in order to speed them up and lower their memory consumption. We then analyze their bottlenecks and propose several techniques which improve their performance even further. To perform evaluation, we have implemented these algorithms inside Desbordante - a science-intensive, high-performance, and open-source data profiling tool developed in C++. Experiments have demonstrated a performance improvement of up to 3x obtained by reimplemented versions, and, with the application of our techniques, up to 10x. Memory consumption has been lowered by up to 2.9x. Subjects: Databases (cs.DB); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Performance (cs.PF) ACMclasses: H.3; I.5; J.0 Cite as: arXiv:2607.23632 [cs.DB] (or arXiv:2607.23632v1 [cs.DB] for this version) https://doi.org/10.48550/arXiv.2607.23632 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: 35th Conference of Open Innovations Association (FRUCT), Tampere, Finland, 2024, pp. 413-424 Related DOI: https://doi.org/10.23919/FRUCT61870.2024.10516381 Focus to learn more DOI(s) linking to related resources Submission history From: George Chernishev [view email] [v1] Sun, 26 Jul 2026 12:36:31 UTC (136 KB)

[AI-85] DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory

链接: https://arxiv.org/abs/2607.23614
作者: Xingyang Yu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); High Energy Physics - Theory (hep-th)
备注: 16 pages, 2 figures, 9 tables. Code, benchmark, and all per-attempt records: this https URL

点击查看摘要

Abstract:We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral-ring proxy. A claim that passes receives a consistency certificate, which states that no tested inconsistency was found, not that the duality is proven. We use the verifier as a repair environment for language-model agents, which receive a deliberately broken claim and must edit it until it certifies. On a preregistered benchmark of 145 broken claims, with the analysis fixed before the first confirmatory model call, verifier-gated retry improves final repair success over a single attempt by +8.3 percentage points (pp) on deepseek-chat and +7.1 pp on qwen-plus (Holm-adjusted p0.002). Under an equal budget of eleven attempts, the stop-first strategy portfolio underperforms independent verifier-filtered resampling by 10.3 percentage points on deepseek-chat but outperforms it by 14.7 points on qwen-plus, reversing the ordering of the two tested verifier-exploitation policies across the two confirmatory models. On qwen-plus, category-level verifier feedback is worth +8.7 pp over content-free retry, and interpretable obligation identities alone are worth +6.4 pp over structurally identical masked feedback. Neither effect is detected on deepseek-chat. Separately, a preregistered MiniMax-M2.5 extension again finds an iteration gain and independent verifier-filtered resampling outperforming the strategy portfolio. Which policy is better thus differs between the two models, while every winning policy uses the same cheap certificate. The verifier, benchmark, protocol, and all per-attempt records are released.

[AI-86] Hybrid Advantage Estimation with Unified Critic for VLM Agent ic Reinforcement Learning ECCV2026

链接: https://arxiv.org/abs/2607.23605
作者: Wenxuan Zhang,Yuhui Wang,Donggang Jia,Xiaoqian Shen,Jian Ding,Ivan Viola,Jürgen Schmidhuber,Mohamed Elhoseiny
类目: Artificial Intelligence (cs.AI)
备注: Accepted by ECCV 2026

点击查看摘要

Abstract:Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end training in agentic environments can improve such multi-turn decision-making abilities, current methods mainly rely on either token-wise optimization over concatenated token trajectories or turn-wise optimization with uniform within-turn credit. In this work, we establish theoretical formulations for the two levels of optimization and derive a hybrid advantage that serves both objectives. Furthermore, with an appropriate choice of discount factor and learning target, we prove that a unified critic model can estimate values for both turn-wise and token-wise. As such, we propose HyGAE, an actor-critic framework that jointly optimizes token- and turn-level objectives with the hybrid advantage and unified critic. We conduct extensive evaluations of HyGAE across five multi-turn decision-making environments, where it achieves an average success rate of 91% and a significant improvement of 10% over other methods. Furthermore, we provide an in-depth analysis showing that the exact analytic form of the hybrid advantage and return is crucial for optimization. Project Page: this https URL.

[AI-87] Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models

链接: https://arxiv.org/abs/2607.23602
作者: Liangyu Li,Qingwen Liu,Mingqing Liu
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 26 pages, 7 figures. Includes supplementary material

点击查看摘要

Abstract:Controllers based on sampling and latent world models assign a predicted terminal cost to each candidate action sequence, choose the minimum, execute its first action block, and replan. This rule can fail even when the terminal cost perfectly and accurately reflects the true task objective in the physical world. Residual prediction error can give an infeasible sequence an anomalously low cost, and a larger proposal pool gives such errors more chances to outrank feasible alternatives. We call this conditional failure proposal overgeneration. In Cube candidate execution audits, increasing the total proposal budget from 72 to 288 reduces the feasibility of selection by minimum latent cost from .375 to .062 for position targets and from .344 to .031 for targets defined by position and yaw, although every larger pool contains a feasible sequence. We introduce Adjacent Set Action Reconstruction (ASAR). Among proposals with low cost, ASAR measures density from standardized early action prefixes and reconstructs a full sequence from an adjacent set with a light anchor from the sequence with minimum cost. On a Carry and Release evaluation set of 75 queries, Kernel ASAR improves event completion success over matching selection by 28.0, 24.0, and 18.7 percentage points under latent cost and by 18.7, 20.0, and 17.3 points under a trajectory reachability cost at 72, 144, and 288 proposals. Analysis of finite proposal pools characterizes selection risk from the lower tail, separation by a related radius support statistic, and sequence containment under an explicit local feasibility condition.

[AI-88] Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents

链接: https://arxiv.org/abs/2607.23586
作者: Zhaoxi Zhang,Xiaomei Zhang
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:

点击查看摘要

Abstract:Long-lived AI agents increasingly evolve after deployment by retaining experience, acquiring skills and tools, revising workflows, delegating work, and moving across task phases. This improves adaptation but creates a distinct authorization problem. Tool-enabled agents can turn model errors and prompt injections into consequential external actions; when evolution occurs under a live grant, the subject exercising that authority or the context in which it acts may no longer match what the user evaluated. Evolution can change both the effects reachable under an old grant and the authority required by the task, which may rise, fall, or become incomparable. Existing tool policies constrain actions but do not determine when a grant survives this change. We formulate authorization continuity: when does an existing grant remain valid, how may active authority change, and what boundary must never move? Our state-bound model fixes a transition envelope and an immutable effect ceiling at grant time. The envelope determines whether the grant survives a mutation; below the ceiling, authority may contract freely and expand only under specified evidence conditions. We distinguish requested from realized effects and prove that, under complete mediation, sound effect abstraction, attenuating delegation, and monitor integrity, mutation cannot amplify protected effects beyond the user-issued ceiling. Agent-produced evidence may allocate authority below the ceiling but cannot raise it. Finally, we map six mutation classes to their authorization consequences. Subjects: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) Cite as: arXiv:2607.23586 [cs.AI] (or arXiv:2607.23586v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.23586 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-89] Verification-Notebook Learning for Source-Aware Multimodal Misinformation Detection

链接: https://arxiv.org/abs/2607.23581
作者: Junyuan Tan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Multimodal misinformation verification is challenging because misleading signals may come from different parts of a post and require different forms of evidence. LVLMs are well suited to this task, but their verification performance often depends on the inference procedure applied to each instance. Existing methods improve this procedure through stronger prompting, retrieval, or deliberation, but rarely retain the verification patterns learned from previous examples. We propose Verification-Notebook Learning (VNL), a non-parametric framework that learns an external verification procedure for a frozen LVLM before inference. VNL builds a compact notebook of decision principles, evidence cues, and recurring pitfalls from prior verification experience. The notebook remains fixed during inference and guides the verification of new examples. Rather than updating model parameters or storing demonstrations, VNL records learned knowledge in an artifact that can be inspected directly. Experiments show that VNL consistently outperforms a range of competitive baselines. Further analyses show that the Verification Notebook improves fine-grained source attribution while remaining compact and interpretable, providing an effective way to accumulate verification knowledge without model training.

[AI-90] An Unofficial FastLAS Tutorial: A Programmers Guide FAST

链接: https://arxiv.org/abs/2607.23557
作者: Fabio Aurelio D’Asaro
类目: Logic in Computer Science (cs.LO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 49 pages. A tutorial on the FastLAS input language. All examples are verified against FastLAS 2.2.0 and clingo 5.8.0; the accompanying task files are available at this https URL

点击查看摘要

Abstract:FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of examples, and it searches for a set of logic program rules (a hypothesis) that explains the examples. These notes are a hands-on introduction to writing FastLAS programs. They are organised as a programmer’s guide: syntax first, then a ladder of worked, numbered examples of increasing difficulty. Every self-contained example here has been run against FastLAS 2.2.0 and shows the tool’s actual output. We keep theory to the minimum needed to write correct programs; throughout, set-off notes flag where FastLAS differs from its sibling system ILASP, and where the two learning algorithms (–opl and --nopl) behave differently. The document is intended as an unofficial tutorial to FastLAS 2.2.0, not as an official language specification.

[AI-91] GTIN: A Unified Framework for Joint Event and Time Prediction in Temporal Graphs

链接: https://arxiv.org/abs/2607.23556
作者: Mohammad Ostadmohammadi,Sepehr Kazemi,Hamid R. Rabiee
类目: ocial and Information Networks (cs.SI); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Temporal graphs are increasingly used to model dynamic systems in diverse domains such as social networks, financial networks, and traffic networks. Predicting both what the next event will be and when it will occur in these systems is crucial for understanding and anticipating complex behaviors, but has not been studied much. To address this gap, we propose a unified mathematical framework capable of capturing varying degrees of complexity across temporal graphs. Our framework is flexible and expressive enough to accommodate a wide range of network structures and temporal dynamics. Building upon this analysis, we introduce our novel approach for jointly predicting the next event and its occurrence time. Empirical evaluations across multiple datasets demonstrate that our method consistently outperforms existing techniques, particularly in scenarios involving irregular event patterns and complex temporal dependencies. These findings highlight the potential of our framework as a robust foundation for future research in temporal event prediction.

[AI-92] Neonatal Hypoxic-ischaemic Encephalopathy Classification from the EEG and HRV Signals Using a Conformer based Masked Autoencoder

链接: https://arxiv.org/abs/2607.23554
作者: Shuwen Yu,William P Marnane,Geraldine B. Boylan,Gordon Lightbody
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Signal Processing (eess.SP)
备注: Paper submits to IEEE Transactions on Neural Networks and Learning Systems

点击查看摘要

Abstract:In this paper, we propose the MAEConformer, a novel self-supervised learning framework that combines the Conformer architecture with the Masked Autoencoder (MAE) paradigm for large-scale representation learning from unlabelled electroencephalography (EEG) and heart rate variability (HRV) signals. By integrating convolutional operations with Transformer-based self-attention, MAEConformer effectively captures both local temporal patterns and long-range contextual dependencies in physiological time series. To enhance reconstruction fidelity and representation quality, a multi-resolution short-time Fourier transform (MR-STFT) loss is incorporated alongside the reconstruction objective, enabling the model to jointly learn temporal and spectral characteristics across multiple scales. Modality-specific EEG and HRV MAEConformer models were pretrained on 6,030h and 4,868h of unlabelled recordings, respectively, and subsequently transferred to expert-annotated downstream tasks. Experimental results demonstrate that the learned representations provide strong transferability and data efficiency. In EEG-based hypoxic ischemic encephalopathy (HIE) severity classification, the pretrained MAE-EEG model achieved test AUCs of 97.19% and 96.56% for binary and four-class classification tasks, respectively, outperforming a range of state-of-the-art supervised and self-supervised baselines. On the HRV-based HIE severity classification task, MAE-HRV achieved a test AUC of 82.42%, surpassing both self-supervised Transformer-based and supervised convolutional baselines. These findings demonstrate the effectiveness of MAEConformer for learning robust and transferable representations across multiple physiological modalities.

[AI-93] ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

链接: https://arxiv.org/abs/2607.23537
作者: Qiao Yan,Yihan Wang,Zhenghao Xing,Jiaqi Xu,Pheng-Ann Heng
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains unclear how vision-language models behave under real-world adverse weather with multi-modal inputs. We argue that a key difficulty lies in degraded environmental observability: under fog, rain, snow, and low illumination, multi-modal observations become unreliable and cross-modally inconsistent, posing challenges to scene understanding, and subsequent decision-making. To study this, we introduce \textbfObsDriveBench, a real-world multi-modal benchmark for adverse-weather autonomous driving. Our benchmark is designed with three capability dimensions: \textbfobservability awareness, \textbfspatial reliability, and \textbfrisk-aware decision-making, enabling fine-grained diagnosis of model behavior under degraded observations. We construct the benchmark through observability meta-annotation, scene description, and capability oriented multiple-choice tasks over synchronized camera, LiDAR, and radar inputs, forming a benchmark with over 14k training and 13k test questions. Experiments reveal consistent performance degradation of existing vision-language models. We further introduce \textbfObsDrive model with normal-weather supervised fine-tuning and adverse-weather reinforcement learning, improving robustness across all three capabilities. The dataset and evaluation code will be released at \hrefthis https URL\textttObsDriveBench.

[AI-94] Mission-Level Runtime Assurance for LLM -Assisted ISR Swarms over a Verification-Aware Fabric

链接: https://arxiv.org/abs/2607.23532
作者: Nikolaos Kekatos,Stylianos Basagiannis,Panagiotis Katsaros,Alexios Lekidis,Tom Nianios
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in contested environments. A growing class of their assurance failures arises not within any single platform but across the swarm: individually-compliant actions compose into a mission-level violation: a prohibited objective split across platforms to evade per-platform lim- its, or a collective budget quietly exceeded. Per-platform guardrails miss these by construction, and contested communications let the violation hide behind lost or delayed evidence. We present a three-tier (platfor- m/squad/mission) compositional runtime-verification framework that de- composes a mission policy into per-agent and cross-agent aspects, aggre- gates per-platform verdicts over a verification-aware messaging fabric, and fuses them with an evidence-aware, two-axis (security x complete- ness) algebra whose provenance names the platforms that jointly trig- gered a violation. Because the fabric makes evidence loss and silence observable, unsupported negative verdicts are downgraded to an explicit unknown rather than reported as mission-wide all-clears. On a simulated ISR mission, an indirect prompt injection that causes real LLM planners to split a prohibited collection task across four platforms is invisible to every per-platform monitor yet detected compositionally with full prove- nance; under an injected fault campaign a best-effort central monitor emits silent false all-clears while the verification-aware fabric emits none

[AI-95] Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis

链接: https://arxiv.org/abs/2607.23524
作者: Xinhao Yao,Yuanzhuo Liu,Changhao Wang,Yunfei Yu,Haoran Tan,Yuyao Zhang,Ruifeng Ren,Minlong Peng,Yong Liu
类目: Artificial Intelligence (cs.AI)
备注: Work in Progress

点击查看摘要

Abstract:Deep search is becoming a core capability of modern agent systems, yet it is typically evaluated solely based on end-to-end answer accuracy. This coupled evaluation paradigm entangles retrieval quality, long-context comprehension, evidence verification, and tool-use decisions, making it difficult to determine whether a model truly knows when and how to delegate information seeking to search. To this end: (1) We formalize this meta-capability as Delegation Intelligence in deep search and decompose it into complementary dimensions-Search Decision-Making (recognizing information insufficiency and deciding whether, when, and how to search) and Information Synthesis and Verification (aggregating evidence from multiple sources, judging source reliability, and synthesizing information under noisy, potentially adversarial conditions). (2) To enable disentangled and reproducible measurement, we develop a controllable synthesis pipeline built on document-grounded reverse engineering. This yields a general recipe for constructing controlled deep-search evaluations rather than a single fixed dataset. (3) As a concrete instantiation, we construct DelegSearchBench, together with a disentangled evaluation protocol that isolates each capability dimension by varying document composition and tool access. (4) Across representative models, we demonstrate that deep-search competence cannot be adequately characterized by final-answer accuracy alone…

[AI-96] Impute On-Demand: Adaptive Correlated Time Series Imputation for Changing Environments

链接: https://arxiv.org/abs/2607.23503
作者: Zhichen Lai,Huan Li,Dalin Zhang,Dong Gong,Lina Yao,Christian S. Jensen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Databases (cs.DB)
备注: 15 pages, 9 figures. Accepted for publication in IEEE Transactions on Knowledge and Data Engineering (IEEE TKDE)

点击查看摘要

Abstract:Internet of Things (IoT) applications generate vast amounts of Correlated Time Series (CTS) data that often contain missing values and require imputation. Existing methods emphasize accuracy but often lack adaptability to changing IoT environments: they are vulnerable to sensor failures, cannot selectively impute only incomplete sensors, and use static architectures that do not adapt to resource availability. To address these limitations, we propose AdaCTSi, an adaptive CTS imputer for changing environments. AdaCTSi combines a One-shot Temporal Convolutional Network with a Learned Time-Sensor Index Table to extract and decouple complex spatio-temporal features into sensor-wise embeddings, enabling adaptation to varying sensor subsets. Sparse Spatial Attention efficiently extracts dynamic spatial correlations, while Correlation-Weighted Sensor Selection selects informative sensors to provide sufficient spatial context. Experiments with twelve baseline methods, three adaptability scenarios, and five benchmark datasets covering traffic, air quality, and trajectory data show that AdaCTSi reduces MAE by an average of 33.1% relative to the strongest baseline on each dataset. A single trained model supports sensor-subset and resource-adaptive inference, and its modest memory footprint enables deployment on commodity computing devices, including MCUs.

[AI-97] Formalizing Flag Algebras in Lean

链接: https://arxiv.org/abs/2607.23500
作者: Gyeongwon Jeong,Seonghun Park,Jihoon Hyun,Sang-il Oum,Hongseok Yang
类目: Logic in Computer Science (cs.LO); Artificial Intelligence (cs.AI); Programming Languages (cs.PL); Combinatorics (math.CO)
备注: 58 pages. Lean code: this https URL

点击查看摘要

Abstract:Razborov’s flag algebra method is a powerful tool for proving asymptotic inequalities in extremal graph theory, often reducing the task to finding a finite certificate by semidefinite programming. We present a machine-checked formalization of the method for finite simple graphs, together with a certificate-to-proof compiler that turns externally generated certificate data into algebraic proofs checked by Lean. The formalization covers the foundations of the method: partially labeled graphs, their densities in large graphs, the quotient algebra of density expressions, graph-limit semantics through positive homomorphisms, and the downward operators used to average out labels. The compiler treats the external semidefinite programming output as candidate data rather than trusted input: Lean independently computes the required density and multiplication facts, verifies positive semidefiniteness exactly over \mathbbQ , and carries out the algebraic normalization steps of flag-algebra proofs. Our case studies yield formal proofs of seven Turán-type upper bounds, including Mantel’s theorem and the Erdős pentagon theorem, a C_4 -density bound for triangle-free graphs, and edge-density bounds for K_4 -free, K_5 -free, and C_5 -free graphs. Independently of the compiler, we formalize the matching constructions that complete the exact Turán densities of Mantel’s theorem and the Erdős pentagon theorem, and prove two inequalities of Goodman. Our constrained semantics also prompted a meta-theoretic comparison of two ways of imposing graph constraints: building a hereditary constraint into the flag algebra from the start, or testing inequalities afterward on constrained graph limits with labels chosen at random. We state the resulting root-plantability criterion characterizing when the two approaches agree; a forthcoming paper will present the complete account.

[AI-98] Do LLM s Know Their Vulnerable Scenarios?

链接: https://arxiv.org/abs/2607.23496
作者: Ziheng Peng,Huiqi Deng,Haoran Jing,Xuankun Rong,Jiahui Han,Xiting Wang,Na Zou,Xia Hu
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 16 pages, 10 Figures, Under Review

点击查看摘要

Abstract:Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear. Meanwhile, mechanistic interpretability studies have characterized both refusal directions and jailbreak-associated features, without explaining the relationship between the two representations. In this work, we show that scenario-wrapped prompts activate internal scenario directions whose causal steering consistently reduces refusal scores. Building on this finding, we propose \textscConcept2Scenario, a concept-based attribution framework for vulnerable scenario discovery. It instantiates a broad concept space with a sparse autoencoder, attributes refusal suppression to individual concepts, translates the identified concepts into interpretable natural-language scenarios, and identifies synergistic scenario combinations through interaction attribution. Across three open-source models, two safety benchmarks, and six black-box jailbreak methods, the discovered scenarios serve as reusable priors that improve average attack success rates by up to 18.2 percentage points. They also transfer to GPT-5, Claude-Haiku-4.5, and Gemini-3-Flash, suggesting that some scenario-level refusal vulnerabilities are shared across model families. Moreover, the identified combinations outperform their individual constituents and enable iterative attacks to succeed in fewer turns.

[AI-99] ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour

链接: https://arxiv.org/abs/2607.23478
作者: Jianhang Xie,Sicheng Tan,Vishnu Naresh Boddeti,Zhichao Lu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Code: this https URL

点击查看摘要

Abstract:Fully homomorphic encryption (FHE) provides strong cryptographic guarantees for private inference, but deploying transformer models under FHE remains prohibitively expensive. A key bottleneck is that non-linear operations such as softmax, normalization, and activation must be replaced with polynomial approximations compatible with the CKKS scheme, and the multiplicative depth consumed by these approximations dominates inference cost. Recent frameworks have advanced approximation techniques, yet all rely on manually configured approximation hyperparameters (e.g., number of iterations, polynomial degree), applied uniformly across all layers. While convenient, this uniform-configuration approach is overly rigid: different layers can tolerate different levels of approximation error without degrading predictive accuracy, and uniform configurations cannot exploit this variability to reduce latency. Allowing each layer to adopt its own configuration, however, causes the search space to explode with model depth, reaching roughly 10^84 configurations for BERT/ViT (12 layers) and 10^225 for LLaMA3 (32 layers), rendering manual exploration practically impossible. We present ATLAS, an automated framework that configures per-layer approximation settings by formulating the problem as a multi-objective optimization over latency and predictive accuracy. The resulting problem is inherently difficult: 1) competing objectives over a large decision space (120 or 320 variables for BERT/ViT or LLaMA3); 2) expensive evaluation, as each configuration takes 70-1,000 seconds even in cleartext; and 3) sparse optimization signals, as 35-50% of candidate configurations yield numerically invalid solutions. ATLAS addresses these challenges through a two-stage optimization strategy that progressively relaxes layer-wise constraints, combined with surrogate models to accelerate evaluation.

[AI-100] Constraint-Bound Agnostic Bayesian Optimization: One Model for All Thresholds

链接: https://arxiv.org/abs/2607.23448
作者: Jin Wang,Xi Lin,Handing Wang
类目: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Expensive constrained optimization problems in real-world industry design often involve constraint thresholds that are difficult to determine in advance. Engineers may need to adjust constraint thresholds to explore different feasibility-performance trade-offs, requiring solutions under a wide range of threshold settings. However, existing constrained Bayesian optimization methods treat each threshold configuration independently, leading to repeated optimization and failing to exploit the shared relationship among continuously varying thresholds. To address this challenge, we propose constraint-bound agnostic Bayesian optimization (CBA-BO), a learning-based framework that learns a parametric constraint model mapping thresholds to optimal solutions. Once learned, CBA-BO directly predicts solutions for arbitrary unseen threshold configurations without additional optimization, with a one-step Bayesian optimization refinement further improving solution quality. Experiments on benchmark and engineering problems demonstrate that CBA-BO learns a transferable threshold-solution mapping, enabling efficient prediction and optimization for arbitrary threshold queries. An intent-guided constraint-bound recommendation mechanism is further developed to improve objective performance while satisfying user-specified constraint preferences.

[AI-101] LA-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation

链接: https://arxiv.org/abs/2607.23425
作者: Arslan Bisharat,Eric Spencer,Brian Ortiz,Khushboo Bhadauria,Mujtaba Nazari,Beatriz Santos,Anisa Ramos,TaiNing Wang,George K. Thiruvathukal,Konstantin Läufer,Mohammed Abuhamad
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 17 pages, appendix included. Introduces TLA±Bench, an execution-grounded benchmark and dataset for natural-language to TLA+^{+} specification generation. Dataset, evaluation code, and model outputs available at publication

点击查看摘要

Abstract:Large language models increasingly write TLA ^+ formal specifications from natural-language descriptions, but progress is hard to measure: existing resources grade by resemblance to a reference or by whether the output parses, neither of which shows correctness. We present TLA ^+ -Bench, a dataset and benchmark that grades by execution. Every gold specification ships a configuration the TLA ^+ model checker runs over the full reachable state space, deciding exactly whether the specification holds the properties that configuration names. The dataset holds 403 model-checked gold and 897 parse-only silver specifications from 13 public repositories, subsumes prior TLA ^+ generation data, and carries four model-written descriptions in two styles from two providers, with difficulty and category labels. Our main finding is about measurement itself: an exact oracle gives not one correctness number but a range. Varying only the grading choices earlier benchmarks leave unstated, on one fixed set of model outputs, the correct rate moves sixfold, from 10.0% to 1.7%; adding the interface-supply choice, where the model is told the configuration’s names, widens the range to elevenfold, from 18.7% to 1.7%. We call this range the correctness envelope and measure each of its bounds. The findings inside it are stable. Every model writes valid TLA ^+ far more often than correct TLA ^+ : the strongest is correct 16% of the time by default and 26% when given the interface names, open models at most 1%, and correctness falls sharply with difficulty.

[AI-102] NeurGO: Learning to Generate Elite Candidates for Meta-Black-Box Expensive Optimization

链接: https://arxiv.org/abs/2607.23408
作者: Jintao He,Huixiang Zhen,Wenyin Gong
类目: Artificial Intelligence (cs.AI)
备注: 18 pages, and 5 figures

点击查看摘要

Abstract:Expensive black-box optimization is ubiquitous in science and engineering, where function evaluations are costly and the evaluation budget is limited. Traditional evolutionary algorithms and Meta-BlackBox Optimization (MetaBBO) approaches typically consume most evaluations on candidate selection, often wasting precious budget on inferior solutions. Although surrogate-assisted evolution and Bayesian optimization aim to reduce evaluations through surrogate models, constructing an accurate global model from limited data remains challenging, and model bias can easily trap the search in local optima. To overcome these limitations, we propose NeurGO, a generative MetaBBO framework that directly synthesizes elite candidates from historical population states. Specifically, we employ an attention-based encoder to capture the population-level search trend and condition a decoder on this representation to generate high-quality candidates, avoiding the expensive evaluation of large offspring pools. We then design a quality-diversity loss to maintain solution quality and population diversity throughout the search. Through extensive benchmarking on CEC 2008 and the COCO BBOB test suites, our method achieves better optimization performance under the same evaluation budget and exhibits faster convergence.

[AI-103] Blood Pressure Estimation from PPG: A Comparative Study of Direct and ECG-Mediated Deep Learning Pipelines

链接: https://arxiv.org/abs/2607.23406
作者: Bo Wu,Haoling Wang,Zhuodiao Kuang,Kateryna Shapovalenko
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Applications (stat.AP)
备注:

点击查看摘要

Abstract:Continuous cuffless blood pressure (BP) monitoring is essential for connected health systems and wearable devices, enabling early detection, longitudinal tracking, and personalized management of cardiovascular disease. Many prior approaches attempt to estimate BP indirectly by reconstructing electrocardiography (ECG) from photoplethysmography (PPG), assuming ECG provides a stronger physiological link to BP. However, ECG sensing is less accessible in wearable settings and may introduce unnecessary complexity. In this work, we first perform a large-scale physiological correlation analysis on the MIMIC-III waveform database, revealing that PPG exhibits substantially stronger coupling with arterial blood pressure (ABP) ( |r|=0.247 , p0.001 ) than ECG does ( r=0.018 , p=0.187 ), challenging the assumption that ECG provides a superior intermediate representation. Motivated by this insight, we conduct a systematic comparison between direct PPG-to-BP prediction and ECG-mediated pipelines using multiple state-of-the-art deep learning models. Across 1.74M segments from 3,127 patients, direct PPG-to-BP prediction achieves British Hypertension Society Grade A performance ( \mathrmMAE_\mathrmSBP = 4.82 mmHg , \mathrmMAE_\mathrmDBP = 4.31 mmHg ), outperforming all ECG-mediated approaches, which achieve only Grade B accuracy. Our findings suggest that accurate continuous BP monitoring can be achieved directly from wearable PPG signals, enabling simpler, more efficient pipelines for real-world connected health systems. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Applications (stat.AP) Cite as: arXiv:2607.23406 [cs.LG] (or arXiv:2607.23406v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.23406 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Kateryna Shapovalenko [view email] [v1] Sun, 26 Jul 2026 01:44:12 UTC (302 KB)

[AI-104] Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning

链接: https://arxiv.org/abs/2607.23394
作者: Adhyyan Narang,Artin Tajdini,Claire Zhang,Jamie Morgenstern
类目: Artificial Intelligence (cs.AI)
备注: 21 pages, 4 figures. Code: this https URL

点击查看摘要

Abstract:Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly. Standard defenses, such as data filtering, mixing in harmless data, and regularization, attenuate these effects but do not eliminate them. We instead pursue robustness through redundancy: collecting multiple datasets from different sources and only learning what is common between them. Thus, if only a subset of sources are malicious, the misbehavior will be blocked. In order to implement this defense strategy, we fine-tune a separate reference model on each source’s dataset and aggregate their next-token distributions at decoding time. We introduce two consensus decoders: a token-wise minimum, which caps each token at the lowest probability any source assigns, and a base-relative variant, which reverts to the base probability on any token the sources move in opposing directions. We further relax exact agreement to tolerate partial support across sources and different surface expressions of the same intention. Across controlled poisoning tasks, subliminal learning, and emergent misalignment, consensus decoding suppresses source-specific misbehavior while preserving shared desirable behavior, including cases where union training and weight averaging retain the unwanted behavior.

[AI-105] Key-Interval A*: Accelerating Grid Pathfinding via Structural Abstraction

链接: https://arxiv.org/abs/2607.23393
作者: Taiquan Sui
类目: Artificial Intelligence (cs.AI)
备注: 9 pages, 1 figure

点击查看摘要

Abstract:Existing exact methods for 4-connected grid pathfinding reduce online search, but often either retain fine-grained search states or require substantial preprocessing. This paper presents Key-Interval A* (KIA*), an optimal pathfinding algorithm that uses lightweight preprocessing to construct and search over a compact interval-level abstraction of free space. KIA* represents free space using intervals: maximal contiguous runs of traversable cells. It extracts key intervals that capture structural boundary changes and connects them through contiguous non-key regions. KIA* then performs A*-style search on the resulting key-interval graph and constructively reconstructs grid paths from interval chains, without cell-level local search. We prove the completeness and optimality of KIA* on 4-connected grids. Experiments on standard benchmarks show that KIA* preserves exact shortest-path lengths and achieves the fastest runtime on seven of eight benchmark groups, with the largest gains on structured and game maps.

[AI-106] Directional Influence Function: Estimating Training Data Influence in Constrained Learning

链接: https://arxiv.org/abs/2607.23388
作者: Xin Wang,R. Tyrrell Rockafellar,Xuegang(Jeff)Ban
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training samples in- fluence the model solution (e.g., learned parameters) is crucial for interpretability and robustness. The classical influence function (IF) estimates sample contribu- tions via local sensitivity analysis, measuring how the solution changes when a specific training sample is perturbed or removed. However, IF becomes unreli- able in constrained settings: data perturbations can reshape both the objective and the feasible region, leading to estimates that violate feasibility. In response, we propose the Directional Influence Function (DIF), a novel estimator that explicitly incorporates these constraints into influence estimation. DIF formulates the opti- mality conditions of constrained learning as a variational inequality (VI) and ana- lyzes how perturbing training data affects this VI. We validate DIF on constrained linear regression and demonstrate that it recovers leave-one-out retraining results, whereas IF and penalty-based IF exhibit significant bias. We further apply DIF to fairness-constrained CNNs, where DIF accurately predicts test loss changes under data removal and aligns closely with actual retraining. Our results establish DIF as an efficient and reliable tool for data attribution in constrained learning.

[AI-107] Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

链接: https://arxiv.org/abs/2607.23386
作者: Paul Simpson,John Kozak,Lisa Doake
类目: Artificial Intelligence (cs.AI)
备注: 43 pages, 9 figures. Working paper v3.13.0. Benchmark data, rule encodings and per-case result artefacts: this https URL

点击查看摘要

Abstract:We document a failure class in frontier large language models – exception chain collapse – observed in eligibility evaluation under nested conditional rules of the form “A is required UNLESS B applies, UNLESS C overrides B”. The failure reproduces at first observation, but its empirical surface is unstable: between March and April 2026 several failure cells closed silently under the same model alias, with no version bump (GPT-5.4 on construction insurance moved from 96.6% to 100%, same prompt and harness). For regulated workflows, frontier-model accuracy is a moving compliance boundary that shifts without notice. We present the Aethis Eligibility Module, a neuro-symbolic architecture in which LLMs author rules from authoritative sources and an SMT-based layer executes them deterministically, consistent with the authored specification regardless of model drift, reasoning-effort defaults, or prompt format. Three evidence bases: (i) a controlled benchmark of 225 scenarios across four regulatory domains documents the pattern and, in replication, the drift that partially closed it; (ii) a 20-scenario adversarial extension on construction insurance, where the engine scores 20/20, as does one of four frontier configurations (GPT-5.4 at low reasoning effort), while the other three, including Anthropic’s strongest model at evaluation time, fail the same coverage-gap edge case; (iii) external validation on nine peer-reviewed LegalBench tasks, 949 held-out cases, where the engine is significantly more accurate than all three frontier models (combined McNemar’s p = 0.003), with margins up to +41 points on the curated multi-prong tasks against the Anthropic models. The contribution is to relocate uncertainty from the inference boundary, where it is silent, to the specification boundary, where it is deliberate and audited. All scenarios, rule encodings, and results are public and reproducible.

[AI-108] Fair Division with Strictly Increasing Valuations: A Tight Threshold for Two-Agent EF1 and PO

链接: https://arxiv.org/abs/2607.23367
作者: Nicholas Teh
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Theoretical Economics (econ.TH)
备注:

点击查看摘要

Abstract:We study whether strictly positive marginal values restore the compatibility of envy-freeness up to one good (EF1) and Pareto optimality (PO) for indivisible goods. For two agents, we identify the exact threshold in the number of goods. Every instance with at most seven goods and strictly increasing valuations admits an allocation that is both EF1 and PO, without any submodularity assumption. In contrast, we construct an eight-good instance with normalized, integer-valued, strictly increasing, submodular valuations in which every EF1 allocation is strictly Pareto dominated. Thus, eight goods are necessary and sufficient for a two-agent counterexample. Finally, we strengthen the three-agent NP-hardness result of Chandramouleeswaran and Nimbhorkar (2026): deciding whether an EF1 and PO allocation exists remains NP-hard for normalized, integer-valued, monotone submodular valuations even when zero marginals are confined to eight fixed agent-good pairs, all involving a single agent.

[AI-109] On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems

链接: https://arxiv.org/abs/2607.23365
作者: Muhammad Tukur,Hayatullahi B. Adeyemo,Tao Chen,Nour Ali,Anis Zarrad,Rick Kazman,Marco Agus,Rami Bahsoon
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 64 pages, 7 figures, 8 tables, submitted to ACM Transactions on Software Engineering and Methodology (TOSEM)

点击查看摘要

Abstract:Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education. While these systems offer powerful data-driven and adaptive capabilities, their complexity, rapid evolution, and dependence on dynamic data pipelines introduce new forms of engineering liability collectively referred to as AI Technical Debts (AITDs). AITDs arise from root causes spanning data governance, model implementation, algorithm design, architectural decisions, operational processes, documentation practices, and testing adequacy. Unlike conventional technical debt, many AITDs are latent and propagate across tightly coupled AI pipelines, leading to maintenance challenges, reliability degradation, and heightened safety or security risks. Guided by the principles of AI Trust, Risk, and Security Management (AI TRiSM), this study reinterprets technical debt through the interconnected dimensions of trustworthiness, focusing on AI safety and security technical debts. We conduct a systematic review of 60 primary studies and identify 31 distinct types of AITD, which are organized into a root-cause-oriented taxonomy comprising seven classes. The analysis examines how these debts map to 18 trust-related concerns, including 6 safety hazards and 12 security vulnerabilities. To support mitigation, the review synthesizes 34 actionable guidelines (8 safety and 26 security) targeting the prevention, detection, and reduction of AITDs across the AI lifecycle. Building on these findings, we introduce AITD-MAP, an integrated framework that connects the AITD taxonomy, quality and risk impacts, and mitigation strategies into a unified structure for risk-aware AI engineering. The framework aims to assist AI software engineers in making AI safety and security technical debts visible, understanding their root causes, and mitigating their presence.

[AI-110] raining with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

链接: https://arxiv.org/abs/2607.23333
作者: Chanwoo Park,Asuman Ozdaglar
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We revisit the regret loss framework introduced in Park et al. (2025), which uses decision-theoretic regret as a direct loss function for training models to make better decisions, through the lens of probability-simplex policies. Our first result shows that a single-layer self-attention model trained with regret loss admits a stationary point whose forward-pass exactly matches smoothed fictitious play with the appropriate stepsize that ensures no-regret behavior-i.e., for any given policy input, the model outputs the same update that smoothed fictitious play would produce. In parallel, we also newly introduce a swap-regret loss function, which extends the regret-loss framework beyond external regret and enables models to directly optimize for swap-deviation robustness. We further show that this swap-regret loss admits a stationary point whose forward pass implements the corresponding swap-regret update induced by classical Blum-Mansour no-pass implementation algorithm, with each head implementing an external-regret update via smoothed fictitious play. Together, these results show that regret-trained attention can realize differentiable mechanisms whose deployment induces equilibrium behavior in games: external-regret dynamics lead to coarse correlated equilibrium, while swap-regret dynamics lead to correlated equilibrium. Thus, regret-based objectives steer minimal attention architectures toward online-learning dynamics with game-theoretic guarantees, without supervised traces of those algorithms.

[AI-111] ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications

链接: https://arxiv.org/abs/2607.23326
作者: Toby Liang,Gopal Sarda,Sagar Davasam,Vikas Yadav
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challenges due to complex system constraints and unexpected user behaviors. Among these applications, slot filling is essential for converting unstructured input into structured, actionable data. In this work, we introduce ESF-Bench, a challenging Enterprise Slot Filling benchmark consisting of 810 multi-turn samples and 6530 slots over 8 unique domains. Curated using a taxonomy of the 57 most challenging slot-filling scenarios observed during real-world enterprise deployments, ESF-Bench exposes notable limitations in current state-of-the-art LLMs, with GPT-OSS-120b low successfully extracting slots for only 20.7% of benchmark samples. To support continued research in this area, we publicly release the benchmark dataset, taxonomy, and accompanying evaluation code on GitHub.

[AI-112] Ordered Network Analysis of Epistemic Emotions during Collaborative Problem Solving

链接: https://arxiv.org/abs/2607.23317
作者: Sifatul Anindho,Videep Venkatesha,Jaclyn Ocumpaugh,Nathaniel Blanchard
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注:

点击查看摘要

Abstract:Investigating how affective states such as confusion and frustration persist and transition during co-situated collaborative problem solving (CPS) is important for understanding the dynamics of epistemic emotions. However, the accurate identification of affective states remain challenging as there is no gold-standard truth in this space. Here, we analyze affective states collected through retrospective cued-recall during an in-person CPS task. Using ordered network analysis (ONA), we examine (1) the overall ordered structure of affective states and how this structure differs across self-caught and probe-caught reporting methods, and (2) what aspects of this ordered structure are emphasized differently in slower and faster groups. We find that ONA reveals differences in persistence and transition patterns that are not apparent from descriptive summaries alone. In particular, we observe a stable epistemic core linking curiosity, optimism, and confusion, with different reporting methods emphasizing different connections among states. An analysis between faster and slower groups show that roles of confusion and disengagement also shift significantly during collaboration, particularly in their relationship to conflict. We interpret our findings in the context of collaboration and discuss their implications in developing AI systems that support CPS.

[AI-113] FILLER: Feature Imputation via Latent Location Exploration and Retrieval

链接: https://arxiv.org/abs/2607.23295
作者: Santu Mondal,Chayan Maitra,Rajat K. De
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In real-world machine learning applications, incomplete observations create a fundamental challenge. Researchers have come up with several ideas to address this crucial problem. However, current models still face challenges in balancing scalability and structural consistency. This study proposes a feature imputation method, called FILLER, that deliberately searches the two-dimensional latent space produced by a generative model and fills the missing values with appropriate entries. The generative model is trained on fully observed data to generate samples from the latent space, and FILLER uses this trained model to impute the values missing in the corrupted test samples. In this study, G-NeuroDAVIS serves the purpose of the generative model. This work also presents a mathematical proof on the convergence of the iterative search. Finally, FILLER has been evaluated on several image datasets under random and structured missingness patterns with varying levels of imputation complexities. In order to justify the efficacy of FILLER, it has been compared against existing state-of-the-art solution strategies in terms of RMSE, PSNR, and SSIM. In addition, Wilcoxon signed-rank test has been carried out to validate statistical significance. Moreover, downstream analyses (classification and clustering) have also established the quality of imputation in terms of standard metrics.

[AI-114] RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning

链接: https://arxiv.org/abs/2607.23290
作者: Xi Chen,Hongru Zhou,Shiyu Feng,Hanyu Zhou,Huahui Yi,Rongsheng Wang,Tiancheng He,Kun Wang,Pingping Liu,Qiankun Li,Sicheng Lin,Huiying Ou,Xiaohong Zheng,Tianying Zang,Zhuohang Wu,Leheng Jiang,Kexin Cao,Wenhan Zhang,ChengYi Li,Zhiyang Wang,Songlin Li,Benyou Wang,Ningbei Yin,Shaoting Zhang,Weili Fu,Jian Li,Kang Li
类目: Artificial Intelligence (cs.AI)
备注: 100 pages 7 figures

点击查看摘要

Abstract:Rare diseases collectively affect an estimated 3.5% to 5.9% of the population, yet more than 70% of patients are misdiagnosed and many endure years of evaluation before a diagnosis is reached, because early presentations are nonspecific and relevant expertise is scarce and unevenly distributed. Artificial intelligence could provide support, but existing systems address isolated stages of care, overwhelmingly diagnosis. They typically depend on the results of downstream investigations, and they treat the variability between models as noise to be eliminated. Here we present RareLens, a system that supports clinical decision-making across the entire rare disease trajectory by exploiting this variability. When heterogeneous large language models evaluate the same case, they generate divergent but complementary reasoning, which RareLens aligns and calibrates into a single convergent, actionable decision at each stage. Four coordinated modules perform primary-visit risk screening, diagnosis, treatment planning and prognosis. Developed and evaluated on RareBench, a real-world dataset of 157,525 cases spanning all 33 Orphanet categories and more than 7,000 conditions, RareLens outperformed every frontier model tested, including GPT-5, DeepSeek-R1, Claude-3.7-Sonnet and Gemini-2.5-Pro, at each stage. It achieved an area under the curve of 0.917 for screening and top-1 accuracies of 65.5% and 89.8% for diagnosis and treatment. In an external study spanning 1,287 cases and 23 physicians, autonomous RareLens and physicians assisted by RareLens both substantially outperformed unaided physicians. These findings indicate that aligning divergent model reasoning, rather than scaling a single model, offers a generalizable strategy for high-uncertainty clinical decision-making.

[AI-115] opoFE: topology-aware LLM -guided Automated Feature Engineering

链接: https://arxiv.org/abs/2607.23286
作者: Sha Li,Naren Ramakrishnan
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective is to discover predictive feature transformations from an exponentially large search space. Recent advances in large language models (LLMs) have expanded the expressiveness of AutoFE by enabling feature program generation beyond predefined operator libraries. However, existing LLM-based approaches remain fundamentally limited by stateless generation and homogeneous search: feature proposals are produced from static prompts without accumulating search experience, while single-population exploration quickly converges to dominant transformation patterns and rarely discovers complementary feature compositions across transformation families. We propose TOPOFE, a topology-aware multi-island evolutionary framework for LLM-guided feature engineering. TOPOFE combines family-specialized exploration, adaptive prompt memory, and topology-guided knowledge transfer to efficiently discover diverse and compositional feature programs. Experiments on 29 public tabular datasets demonstrate consistent improvements over state-of-the-art AutoFE methods across classification and regression tasks. Beyond predictive performance, TOPOFE discovers more diverse and transferable feature programs that generalize across multiple downstream predictors and LLM backbones.

[AI-116] X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference

链接: https://arxiv.org/abs/2607.23264
作者: Jianwen Xian,Zhiyuan Xu,Yuchen Li,Ziliang Lai,Kang He,Zhen Huang,Aichen Feng,Jinyan Chen,Yilin Zhang,Qinqin Chen,Chengru Song
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Tensor Core computation. Existing systems schedule when communication is issued and when received data becomes consumable, but omit post-issue progress before remote-visible completion, making sender backpressure hard to predict. We identify X-Stage, a software-visible post-issue pipeline stage. Measurements on an eight-GPU node with a recent NVIDIA architecture show that short remote-store bursts drain as the issuer resumes work, whereas sustained injection exhausts finite outstanding capacity and delays later issues. A lightweight Burst-Gap model parameterized by backpressure-free issue time, effective drain rate, and outstanding capacity predicts issue overhead, recovery between bursts, and the onset of backpressure. Guided by the model, we redesign two communication-computation fused kernels. For DeepGEMM MegaMoE, interleaving Linear-1 and Linear-2 work across expert waves places computation between concentrated remote-store bursts, yielding a 1.18x geometric-mean and 1.62x maximum kernel speedup over the Expert-Wave baseline across 84 configurations. For Ulysses sequence-parallel attention, tile-granular fusion of the post-attention All-to-All with FlashAttention lets an output-tile owner issue remote stores and resume computation without a dedicated communication warp or streaming multiprocessor. FlashAttention-3 and FlashAttention-4 reach maximum sender-visible speedups of 1.43x and 1.42x over serial execution, and at long sequences their steady-state times approach those of FlashAttention alone. These results establish post-issue progress as a measurable scheduling lever for shaping bursts, avoiding backpressure, and hiding sender-side overhead. Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.23264 [cs.DC] (or arXiv:2607.23264v1 [cs.DC] for this version) https://doi.org/10.48550/arXiv.2607.23264 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-117] SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

链接: https://arxiv.org/abs/2607.23263
作者: Yang Wan,Zhenhao Zhang,Jierui Wang,Linchao Zhu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Deciding whether a trajectory actually fulfills its instruction governs how we measure computer-use agents on long-horizon graphical-user-interface tasks and how we train them with reinforcement learning. This judgment has long relied on rule-based evaluation, which struggles to align with human intention and goes stale when an app updates or its online content drifts. Existing model-based judges attempt to address these problems but still leave a performance gap to the rule-based evaluation. We propose the \textbfSeekJudge framework, in which four role-specialized agents, a Condense, a Ground, a Seek and an Analyze agent, reach a verdict through a Seek–Analyze loop over the trajectory. A seed-calibrated distillation pipeline trains one specialized 9 B model to serve as the shared backbone for all four agents. Measured by downstream success rate on held-out RL test goals, SeekJudge is the first practical model-based reward to match or surpass native rule-based supervision in online RL. Beyond accuracy, SeekJudge provides step-level judgments, runs far cheaper than a closed-source large model, and keeps a small per-call context that scales to much longer trajectories. We further contribute a general architectural improvement to the reward server that speeds up judging in RL. Together these make model-based reward a practical drop-in for rule-based supervision in CUA reinforcement learning.

[AI-118] CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics

链接: https://arxiv.org/abs/2607.23258
作者: Xinhong Xu,Yimeng Zhang,Yuanlong Zhang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains unresolved: \textbfwhether a model pretrained on one collection of recordings can generalize to new datasets, experimental paradigms, and even species. Existing approaches are often designed for specific tasks and evaluated on a single dataset, making it unclear whether their learned representations are reusable for new calcium trace datasets. To tackle this gap, we present \textbfCAPT, a \textbfContinuous \textbfAutoregressive \textbfPopulation \textbfTransformer for calcium population dynamics. CAPT models continuous calcium traces directly through a continuous patch tokenization strategy and is trained autoregressively, enabling end-to-end pretraining and adaptation to diverse downstream tasks. We first pretrain CAPT on a large-scale mouse calcium imaging dataset and evaluate its transferability across independent mouse, larval zebrafish, and \textitC. elegans datasets collected by different laboratories. In these transfer settings, the pretrained backbone is frozen and only adaptation modules are updated. Across neural population forecasting and behavior decoding tasks, CAPT consistently outperforms specialized and general-purpose baselines. Alongside predictive performance, multimodal analyses using NeuroPAL annotations in \textitC. elegans datasets show that CAPT embeddings form a shared functional space across datasets and capture anatomical cell-identity-related structure. These results suggest that the continuous autoregressive modeling opens up possibilities for a simple route towards general-purpose neural foundation models for calcium imaging, which can generalize across datasets, experimental paradigms, and species.

[AI-119] Characterisation of Density-based FM generation methods in the context of Information Fusion

链接: https://arxiv.org/abs/2607.23243
作者: Yanhao Huang,Christian Wagner
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Fuzzy Integral (FI) based aggregation provides a powerful mechanism for nuanced aggregation, for example, in ensemble approaches or decision-level fusion more generally. The main challenge of this approach is the appropriate parametrization of the Fuzzy Measure (FM), which captures the worths of the individual components–and their combinations–which are being fused. Here, widely used approaches including the Sugeno- \lambda and Decomposable FMs, parametrize the FM by extrapolating from the densities, i.e. the weights associated with individual sources, while respecting the FM’s monotonicity constraint. This paper articulates that this information is, in general, insufficient to uniquely identify a discrete FM; but shows how an interval-valued FM can indeed be determined uniquely. We proceed to show how the incorporation of additional information beyond the above, such as the choice of a specific FI and a dataset, then allows for obtaining even more specific interval-valued FMs. In practice, establishing the quality of an empirically determined FM is not trivial. To help address this, we show how the likelihood with which a resulting interval FM encompasses the ideal', i.e. the commonly intangible, best, or ground-truth numeric FM, can be determined, producing a confidence interval at a given confidence level. Finally, based on a series of experiments, we demonstrate empirically that the Choquet FI output based on this FM can also be regarded as the confidence interval for the ideal’ information fusion result, providing a novel means to characterize FI fusion outcomes a priori and charting a pathway for future research.

[AI-120] Context-Aware Concept Distillation for Trustworthy Flood Prediction IJCAI2026

链接: https://arxiv.org/abs/2607.23237
作者: Eli Levinkopf,Efrat Morin,Claudia V. Goldman
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: to be published in IJCAI 2026 proceedings

点击查看摘要

Abstract:Effective flood risk management relies on accurate forecasting, yet the “black box” nature of stateof-the-art Deep Learning models creates a barrier to trust and accountability in high-stakes public safety decisions. While existing Explainable AI (XAI) methods offer local attributions, they fail to provide the verifiable, operationally meaningful causal narratives required by disaster response authorities. To address this societal challenge, we propose Context-Aware Concept Distillation (CACD), a framework developed in collaboration with domain experts to distill opaque LSTMs into interpretable, hydrology-aware surrogate models. We introduce an unsupervised pipeline to discover a “Hydrological Language” and a Residual Hypernetwork that dynamically modulates these concepts based on static basin characteristics. Evaluated on 5,203 basins globally, our model achieves high fidelity (Median NSE 0.70), significantly outperforming black-box baselines (e.g., Multi Layer Perceptrons) on unseen future data. By demonstrating that human-interpretable concepts are sufficient to reconstruct flood dynamics, this work balances AI accuracy with the transparency required for responsible environmental decision-making.

[AI-121] An Ontology for Machine Learning Interatomic Potentials

链接: https://arxiv.org/abs/2607.23219
作者: Daniel Hernández,Jong Hyun Jung,Yuji Ikeda,Yongliang Ou,Pranav Kumar,Tom Schächtel,Wenchuan Liu,Xin Li,Xi Zhang,Xiang Xu,Lifang Zhu,Fritz Körmann,Steffen Staab,Blazej Grabowski
类目: Artificial Intelligence (cs.AI); Databases (cs.DB); Digital Libraries (cs.DL)
备注:

点击查看摘要

Abstract:Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces—conventionally computed by density functional theory (DFT) or wave-function methods—at a fraction of the cost. The field encompasses a growing ecosystem of algorithms, training datasets, hyperparameters, and target materials, yet the metadata needed to systematically compare, reproduce, and build upon MLIP studies remains scattered across papers, scripts, and ad-hoc file formats. We present the MLIPs ontology, an OWL 2 DL ontology that captures the concepts needed to describe MLIP methods, their hyperparameters, training datasets with DFT provenance, and published benchmarks. The ontology is organized into three modules—Method, Training Data, and Benchmark—and connects existing ontologies in materials science (MDO, CMSO/ASMO) and machine learning (ML-Schema), complementing dataset-side schemas such as Croissant. It declares 27 formal axioms enforcing data completeness and consistency, including property chains that link trained models to their methods and training data. We demonstrate the ontology through a running example based on Moment Tensor Potentials and evaluate it through competency-question execution on a 20-paper seeded knowledge graph, OWL reasoning, and comparison with existing ontologies.

[AI-122] False Prophets: On the Security of World Models in Agent ic Systems

链接: https://arxiv.org/abs/2607.23147
作者: Erik Imgrund,Anna Wimbauer,Klim Kireev,Konrad Rieck
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent to predict the results of its actions. Recent research proposes to enhance predictive capabilities via specially trained environment simulators-world models. While world models can improve performance, they can also mislead agents into executing harmful actions, creating significant security and privacy risks. In this paper, we raise security concerns regarding the usage of world models in agentic systems. We discover a range of world model specific vulnerabilities, which can be exploited in terminal-based agents to execute malicious code or extract sensitive data. To facilitate future development, we introduce a security benchmark dataset designed for text-based world models. We argue that some risks are intrinsic to approximate world modeling, and show that attackers can induce mispredictions in agentic pipelines with up to 95% success rate, possibly resulting in unintended command execution, denial of service, drainage of wallet and private information extraction. Finally, we provide practical recommendations for practitioners to mitigate the discovered harms and harden agentic systems.

[AI-123] SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

链接: https://arxiv.org/abs/2607.23123
作者: Summer Sun(Shaqiu Community)
类目: Artificial Intelligence (cs.AI)
备注: 17 pages, 8 figures. Code and aggregate results: this https URL

点击查看摘要

Abstract:Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation. We introduce SQBench, a benchmark for evaluating production-oriented task delivery by language-model agents. SQBench v1.0 contains 220 standardized tasks organized into L1 atomic capabilities, L2 composite skills, and L3 business scenarios. Each task requires an agent to process input assets, use available tools, and produce an explicitly specified deliverable. The evaluation first computes functional Completion and then derives Risk Penalty and Performance from independently evidenced triggers in a 10D Risk Matrix. A Strict Pass requires Completion = 1 and Risk Penalty = 0. We evaluate 27 model configurations under a common protocol, with one run per configuration-task pair. The highest prespecified Weighted Pass@1 is 60.5%. Mean Strict Pass@1 on L3 is 18.5%, and every configuration performs worse on L3 than on both L1 and L2, indicating that delivery under domain constraints is a shared weakness within the current task set. Of 2,348 results with Completion = 1, 113 (4.8%) fail the Strict Pass criterion because of risks such as unverifiable citations, inappropriate resource use, or format violations. These results show that functional completion alone does not fully characterize delivery quality and that risk determinations should be reported separately.

[AI-124] A scalable online machine learning approach for Stock Recommendation

链接: https://arxiv.org/abs/2607.23120
作者: Harsh Nagarkar
类目: Computational Engineering, Finance, and Science (cs.CE); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC); Computer Science and Game Theory (cs.GT)
备注: 6 pages, 4 figures

点击查看摘要

Abstract:Stock recommendation systems face the dual challenge of adapting to rapidly changing market conditions while maintaining low-latency predictions for end users. Traditional batch-trained models fail to capture concept drift, and monolithic architectures struggle to provide fault tolerance under load. This paper presents a scalable online deep learning-based stock recommendation system built on a distributed microservices architecture using Kubernetes, Docker, and RabbitMQ. The system employs a hybrid leader-follower architecture where a primary model continuously trains on streaming financial data, including EPS, MACD, and price, from the Alpha Vantage API while multiple replica models serve user-facing recommendations in parallel. A multilayer perceptron implemented with TensorFlow Recommenders generates content-based recommendations using explicit user ratings (1-5) and transfer learning. The architecture ensures high availability. The leader persists model weights to Google Cloud Object Storage, allowing replicas to recover seamlessly upon failure, while RabbitMQ provides message durability and replay. Results demonstrate that the system serves stock recommendations in 23 seconds per request and processes up to 500 portfolio addition requests per second per follower. Key limitations include data staleness (up to 150 minutes due to API rate limits) and the absence of a service mesh for inter-cluster security. This work contributes a production-ready reference architecture for online recommender systems that balances consistency, availability, and scalability in a financial domain context

[AI-125] Compiler-Grounded Hierarchical Diagnosis for LLM -Based Triton Kernel Optimization

链接: https://arxiv.org/abs/2607.23089
作者: Dongjie Chen,Ping Zhao,Bohua Zhan,Yulong Wang,Shushu Chen,Liangjun Feng,Hao Zhou,Min Shen,Linmu Wang,Weijia Sheng,Xiangyu Wei,Weijie Ding,Jianhui Huang,Yaoqing Gao
类目: Artificial Intelligence (cs.AI); Programming Languages (cs.PL)
备注:

点击查看摘要

Abstract:Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rely on surface signals such as compilation feedback and profiling metrics. These signals reveal that a kernel is slow, but not why the backend compiler fails to realize a profitable optimization, especially on emerging accelerators such as NPUs. We therefore formulate kernel optimization as a progressive cross-layer diagnosis problem that links runtime symptoms to IR structure and compiler behavior before rewriting source. Based on this insight, we present our system, a compiler-grounded and hierarchical optimization framework for Triton kernels. the system escalates from lightweight pattern triage and profiling diagnosis to IR attribution and compiler-grounded analysis only when deeper evidence is needed, then proposes evidence-backed source-level rewrites. We implement the system on Triton for Ascend NPUs and evaluate it on 37 successfully converted entries from a standardized NPUKernelBench-derived Ascend 950 benchmark. Across these entries, the system attains a geometric-mean speedup of 4.35 \times and a median speedup of 2.73 \times from the initial to optimized Triton kernel; 22/37 exceed 2 \times and 13/37 exceed 5 \times . The complete distribution ranges from near-baseline entries to large wins, motivating transparent reporting of the current system’s scope and limitations. Subjects: Artificial Intelligence (cs.AI); Programming Languages (cs.PL) Cite as: arXiv:2607.23089 [cs.AI] (or arXiv:2607.23089v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.23089 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-126] Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios

链接: https://arxiv.org/abs/2607.23088
作者: Lixun Ma,Ruolong Ma,Bei Wang,Feng Wei,Zhenguang Liu,Lorenzo Cavallaro,Wentao Chen
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) are widely used for code generation, yet their security behavior in realistic development workflows remains underexplored. Existing benchmarks often rely on explicitly specified security requirements, failing to capture real-world scenarios where prompts are frequently ambiguous or incomplete. In this paper, we adopt a developer-centric perspective and identify three representative risk scenarios that commonly lead to security vulnerabilities in LLM-generated code: Ambiguous Requirements, Under-Specified Operational Context, and Security–Functionality Conflict. Based on these scenarios, we construct a large-scale benchmark comprising 2,700 test cases, enabling fine-grained evaluation of LLM security under realistic conditions. Extensive evaluation of eight state-of-the-art LLMs reveals that all models exhibit average vulnerability rates exceeding 56% across risk scenarios. We further demonstrate that security-aware prompting can substantially mitigate these risks, achieving up to 45% improvement.

[AI-127] Scoping Review of AI Metrology and ESG in the Semiconductor Sector: Implications for Safe and Sustainable by Design (SSbD)

链接: https://arxiv.org/abs/2607.23082
作者: Karen Ang,Han-Teng Liao
类目: ocial and Information Networks (cs.SI); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Systems and Control (eess.SY)
备注: to be published in the 32nd IEEE ICE/ITMC Conference (Porto, Portugal) proceedings

点击查看摘要

Abstract:The semiconductor sector faces a dual transition: scaling manufacturing execution through Artificial Intelligence (AI) while satisfying stringent sustainability mandates, such as the EU Carbon Border Adjustment Mechanism (CBAM). This paper presents a scoping review of 1,465 documents indexed in Web of Science and Scopus, spanning AI-integrated metrology, supply chain ESG, and federated industrial data spaces. Network analysis reveals a highly fragmented “core-periphery” knowledge structure, emphasizing a critical structural hole between AI-driven process optimization and downstream sustainability governance. To close these gaps, this study proposes a 6-layer Safe and Sustainable by Design (SSbD) architecture grounded in a System of Systems (SoS) paradigm. By establishing distinct “grid-to-core” and “standards-through-supply-chain” integration pathways, the proposed framework demonstrates how virtual metrology (VM), localized federated learning, and defensive RegTech mechanisms can build provenance-aware data fabrics. Ultimately, this architecture positions regulatory compliance as a driver for innovation, enabling secure, climate-neutral, and circular value chains in semiconductor manufacturing.

[AI-128] Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

链接: https://arxiv.org/abs/2607.23077
作者: Narthana Sivalingam,Santhirarajah Sivasthigan,Buddhi Wijenayake,Roshan Godaliyadda,Vijitha Herath,Parakrama Ekanayake
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approaches perform well but often rely on deep stacks of layers to learn these heterogeneous dependencies implicitly, increasing computational cost. We revisit this problem from a structural perspective and decompose multi-entity temporal dynamics into three interaction types: spatial interactions among entities, temporal interactions across time, and cross interactions coupling the two domains. We propose a structured spatio-temporal transformer block that explicitly models all three within a single stage. It uses parallel spatial and temporal self-attention, followed by bidirectional cross-attention, and combines the outputs through learnable gated fusion. By directly encoding these complementary views, the model reduces the need for deep stacking. We evaluate the approach on video-based group activity recognition, skeleton-based human interaction analysis, and wearable sensor-based activity recognition. Despite its simplicity, the single structured Transformer block matches or outperforms deeper architectures with only 1.76M parameters. The results suggest that depth in prior models partly compensates for implicit and entangled interaction modeling, whereas explicit factorization offers a more efficient and transparent alternative. More broadly, this work supports a structure-first design principle: expressive multi-entity temporal reasoning can emerge by exposing interaction structure rather than relying on depth.

[AI-129] raceable LLM Reasoning for Fake-Order Fraud Detection

链接: https://arxiv.org/abs/2607.23075
作者: Siqi You,Bingsong Xu,Zhixian Zheng,Xinjian Peng,Yang Xie,Ying Wang,Jiarong Xu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 15 pages, 3 figures

点击查看摘要

Abstract:Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approaches often rely on expert-designed features, produce black-box decisions, and provide limited interpretability. To address these limitations, we propose DeepScrub, a reinforcement learning framework built upon large language models (LLMs) for fake-order fraud detection with traceable reasoning. DeepScrub introduces three innovations. First, a semantic unification module converts heterogeneous risk signals into textual descriptions that LLMs can understand. Second, continued pre-training on risk-control corpora injects domain knowledge, and task rewards jointly evaluate prediction correctness and reasoning quality. Third, the SUggest-REflect (SURE) mechanism incorporates expert feedback and model self-checking to iteratively refine reasoning paths. On a real-world fake-order fraud detection dataset, DeepScrub achieves a macro-F1 score of 85.3%, outperforming the best baseline by 2.7 percentage points. Our task-optimized 8B model further surpasses a 32B model, showing that domain adaptation can matter more than model scale in this setting. In a four-week live pilot, DeepScrub achieved 91.8% precision and 88.5% recall, improving over first-stage human reviewers by 16.6 and 38.8 percentage points. It reduced first-stage manual review workload by 94% and saved nearly one million RMB annually. These results show that DeepScrub improves fraud review accuracy, reduces first-stage review workload, and provides traceable evidence for production risk-review workflows.

[AI-130] SymStep: Symbolic Step Verification for Logical Reasoning

链接: https://arxiv.org/abs/2607.23055
作者: Aida Usmanova,Rui Gao,Dilshod Azizov,Ricardo Usbeck,Zangir Iklassov
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps. We introduce SymStep: an LLM makes one atomic claim at a time (DEDUCE: Alice, pet, Cat), then a lightweight constraint propagator checks the claim for consistency with prior accepted deductions, rejects contradictions, and cascades implied facts automatically. SymStep+G additionally provides MRV guidance after each accepted step, directing the LLM toward the most constrained unresolved variable. On a 35-puzzle retained subset of ZebraLogicBench, a benchmark of 1,000 Einstein-style logic puzzles, Direct and CoT both achieve 0%, while SymStep+G reaches 97%. On AR-LSAT analytical reasoning problems, SymStep achieves 100% vs. CoT’s 87%. On LGP-14, SymStep+G achieves 100% vs. 0% for CoT and Logic-LM, the strongest prior symbolic+LLM baseline we compare against. Ablation studies reveal that MRV guidance is a key mechanism for reducing directionless cycling, while consistency checking provides a safety net against explicit contradictions. Across six benchmarks spanning five task domains, SymStep variants match or exceed every baseline on constraint-dense and arithmetic tasks. Experiments on AQUA-RAT algebra confirm the advantage is constraint-density-specific.

[AI-131] MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models

链接: https://arxiv.org/abs/2607.23047
作者: Ashitabh Misra,Madhav Agrawal,Arham Jain,Tarek Abdelzaher
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Preprint

点击查看摘要

Abstract:Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget. In practice the budget varies across deployments and is unknown at calibration time. Adaptive quantization addresses this with one offline calibration that serves any budget, yet current methods score layer sensitivity in a manner that does not consider its dependency on quantization levels of other layers. We show that a layer’s sensitivity depends strongly on the bitwidths of its upstream layers and that this dependence shifts the resulting preferred bit allocation. We propose MixQuant, a technique-agnostic adaptive framework that wraps any base quantizer. MixQuant marginalizes each layer’s distortion over random quantized upstream configurations to obtain budget-agnostic scores, calibrates the quantizer’s parameters on plans the allocator itself produces, and penalizes allocations that leave layers at the lowest bitwidths. A single greedy pass then serves any budget at deployment. Across Llama-3.2-3B, Llama-2-7B, and Mistral-7B under AWQ and GPTQ, MixQuant outperforms adaptive and mixed-precision baselines in every setting, improving average accuracy by up to 8 points and reducing perplexity from 12.43 to 10.70 at the tightest budget, while matching an ILP solver at negligible deployment cost.

[AI-132] Stress-testing large language model agents in a robotic chemistry laboratory

链接: https://arxiv.org/abs/2607.23045
作者: Lulu Guo,Yingkai Sun,Xiaobo Li,Luyao Ge,Ziming Wang,Haitao Zheng,Jingyu Li,Huijuan Zhang,Bingxu Chen,Daobin Liu,Yuebo Liu,Jie Li,Xiaohui Li,Linjiang Chen,Yi Luo,Jun Jiang
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. Here, we use a robotic chemistry laboratory as a physical-world testbed to make scientific agency measurable. Its 45 modular workstations exposed as machine-readable skills enabled 4,608 trials. Only 3.3% of trials produced expert-assessed executable workflows under laboratory constraints; even the best system achieved 28.1%. Long-horizon planning remained a challenge: only three executable workflows exceeded 30 operations, although the longest contained 44. Across five rounds, experimental feedback prompted local adjustments but no workflow-level replanning or analytical-method redesign. By making physical executability and evidence-driven replanning measurable, our study provides an evidence-based assessment of deployment readiness and a diagnostic framework to guide closed-loop improvements towards physically grounded autonomous research.

[AI-133] Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View

链接: https://arxiv.org/abs/2607.23029
作者: Kun Zhao,Xu Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Federated learning enables collaborative model training across distributed clients without centralising their data, yet privacy remains a persistent concern because the shared model updates can leak information about local datasets. Existing privacy-preserving methods either inject calibrated noise into client updates, limiting their composition guarantees, or formulate client privacy choices as a multi-agent game whose Nash equilibrium becomes intractable as the number of clients grows. We bridge these two lines of work by formulating privacy-preserving federated learning as a mean-field privacy game: each client strategically chooses its own privacy budget while interacting with the population only through a single mean-field statistic. The mean-field limit yields a tractable equilibrium for arbitrarily many clients, accommodates heterogeneous client preferences, and inherits an exponentially decaying privacy guarantee through a log-Sobolev contraction. The framework recovers the entropic privacy baseline as the homogeneous special case and the multi-agent privacy game as the finite-population case. Experiments on quadratic regression, logistic regression, and MNIST demonstrate that the proposed framework attains the privacy-utility trade-off of the entropic baseline while delivering a personalized privacy guarantee that the homogeneous baseline cannot express.

[AI-134] All in One: Generative Modeling as Mean-Field Game Design

链接: https://arxiv.org/abs/2607.23026
作者: Kun Zhao,Xu Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Mean-field games (MFGs) offer a unifying lens on continuous-time generative modeling: a cost tuple recovering twelve prominent models—Continuous Normalizing Flows, OT-Flow, Score-based Models, Schrödinger Bridges, and more—as special cases of one variational problem. Yet two dimensions of this space remain entirely unexplored: the interaction term \mathcalI is set to zero in many existing models, and the rich family of MFG solvers has never been applied to generative modeling. We address both gaps with MFGLab an open-source PyTorch library whose primary API is the cost tuple: all twelve models are specified by four composable cost functions, and the training loop, log-Jacobian, and reverse-ODE sampler are shared automatically. We additionally propose DI-Flow, a novel cost design that uses a differentiable entropy functional to encourage mode coverage, and provide learning-based MFG solvers that substantially outperform neural training on stochastic-dynamics rows. Experiments on two 2-D benchmarks confirm that the unified API is lossless relative to hand-coded implementations.

[AI-135] Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming

链接: https://arxiv.org/abs/2607.23019
作者: Zirong Chen,Meiyi Ma
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO); Symbolic Computation (cs.SC)
备注: Accepted at the 20th Conference on Neurosymbolic Learning and Reasoning

点击查看摘要

Abstract:Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound. We present Reason Popper-ly, a neurosymbolic framework that uses inductive logic programming (ILP) to learn relation composition rules from reasoning traces and deploys them as an online verifier for step-level correction. Given an LLM-generated trace, the method checks each inferred step against the learned rule table, diagnoses the violation type, rewrites incorrect steps with symbolically derived repairs, and regenerates the remaining suffix so that the model can produce its final answer conditioned on a verified trace. We evaluate on CLUTRR, a multi-hop kinship reasoning benchmark, using five language models over reasoning chains of 2 to 10 hops. Across all models, Reason Popper-ly consistently improves terminal accuracy over standard CoT, with gains of up to 48 percentage points for small models and 15 points for frontier models on the longest chains. Compared with a fully exogenous symbolic pipeline, our method performs better on harder instances by preserving the model’s successful grounding while correcting only verifiable reasoning failures. In addition, step-level ILP verification yields a fine-grained error taxonomy that provides diagnostic insight beyond final-answer accuracy.

[AI-136] Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop

链接: https://arxiv.org/abs/2607.23002
作者: Jeff Otterson(W. P. Carey School of Business, Arizona State University)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 26 pages. Two pre-registered experiments; protocols, all run receipts, and analysis code at this https URL

点击查看摘要

Abstract:Large language models increasingly write both code and the tests meant to check it; coverage records what ran, not what was verified. We study an adversarial test-hardening loop under a mechanical oracle: a Tester model writes tests, mutation testing names surviving injected defects, and a Critic model writes tests to kill exactly those, with every verdict decided mechanically, so no model judges another’s output. In Experiment 1, on five Python subjects (one same-lineage-loop cell could not be scored), the loop killed 105 mutants that one-shot generation missed and lost none, and the cross-lineage-Critic question returned a pre-declared null. The central finding was an autopsy: an earlier analysis reported a cross-lineage effect at p = 9.5e-66 that was an instrument artifact, an output cap silently truncating the verbose model, caught only by adversarial review of the completed analysis. Review then found a further confound, each arm resampling its own initial suite; Experiment 2 removes it. Under a pre-registered frozen-shared-round-0 design (five replicates on each of four subjects, seeds committed in advance), same-lineage Critic rounds killed 78% of the survivors the frozen initial suite left standing (mean incremental kill rate 0.783, 95% cluster-bootstrap interval [0.592, 0.935]), a within-replicate causal estimate; the cross-provider configuration showed a positive pilot difference (rate gap 0.178, 95% interval [0.039, 0.347]; magnitude dominated by a single replicate) at 5.5x lower arm cost. This compares two named model-provider-harness configurations, not an isolated lineage effect: part of the gap is one configuration’s receipted operational failures, including truncation recurrences, now detected and scored rather than laundered. Cross-model comparisons can inherit the asymmetries of the harness that runs them. We release both protocols, all receipts, and the analysis code.

[AI-137] Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

链接: https://arxiv.org/abs/2607.22997
作者: Qing Yang,Xun Wang,Ziguan Wang,Zhenjiang Li,Hongqiang Wang,Dongdong Weng
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Graphics (cs.GR); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Physical AI – the integration of large vision-language-action (VLA) models with embodied agents that act in the real world – has emerged as the next major frontier for AI, echoed by industry leaders such as Jensen Huang (``the next big thing is Physical AI, AI with a body,‘’ GTC Paris, June 2025) and Dr. Lisa Su (we're entering the world of Physical AI ... this is where AI enters the real world,' CES 2026). This paper presents an end-to-end, fully AMD-accelerated technology stack for embodied manipulation, spanning data-center training silicon, Radeon PRO simulation/rendering GPUs, and Ryzen AI edge compute, unified by the open ROCm software stack. We demonstrate that training and deploying VLA-based manipulation policies does not require a CUDA-locked ecosystem. Four progressive demonstrations are presented: (1) a Sim-to-Real manipulation pipeline trained with SmolVLA and deployed on a physical Franka arm; (2) a semantic, language-grounded object-selection task (one-of-three’); (3) a Real2Sim synthetic-data generation pipeline that fuses 3D Gaussian Splatting (3DGS) reconstructions of real scenes with the Genesis physics engine; and (4) large-scale reinforcement learning for quadruped and humanoid locomotion benchmarked across multiple hardware platforms. All pipelines run natively on ROCm + PyTorch on RDNA4 (Radeon AI PRO R9700) and RDNA3.5 (Radeon PRO W7900) hardware and are reproducible on the free Radeon Cloud Platform.

[AI-138] Label-free Industrial Fault Detection via Adversarial Inverse Reinforcement Learning: A System for Run-to-Failure Prognostics

链接: https://arxiv.org/abs/2607.22987
作者: Dhiraj Neupane,Mohamed Reda Bouadjenek,Richard Dazeley,Sunil Aryal
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Machinery fault detection (MFD) remains heavily reliant on supervised learning, which struggles with the scarcity of fault labels in real-world settings. While reinforcement learning (RL) offers a framework to model the sequential nature of degradation, current ``RL-based’’ MFD methods reduce the problem to a static contextual bandit (CB) formulation: by ignoring state transitions and discarding the temporal discount factor, they collapse to standard supervised classification. We propose an adversarial inverse reinforcement learning (AIRL) framework that treats MFD as an offline IRL problem. Unlike reconstruction-based approaches that rely on static error margins, or CBs that ignore dynamics, our method recovers an intrinsic “health” reward directly from observational state transitions, requiring neither manual reward engineering nor fault labels. On three run-to-failure benchmarks (HUMS2023, IMS, XJTU-SY), AIRL is the only method achieving non-saturated post-detection consistency across all datasets, while CB baselines fail to detect gradual degradation and reconstruction models collapse into always-anomalous states. Code and data: this https URL.

[AI-139] Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI

链接: https://arxiv.org/abs/2607.22953
作者: Sourena Khanzadeh,Daniel Platnick,Marjan Alirezaie,Hossein Rahnama
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Computers and Society (cs.CY); Social and Information Networks (cs.SI)
备注:

点击查看摘要

Abstract:Modern AI systems bring societal risks such as mass surveillance, extreme concentrations of power, and loss of user autonomy—calling into question a model where third-parties collect and control massive amounts of user data. Users require a sovereign system to securely own, govern, and disclose their context while remaining compliant across regulated domains with strict provenance, interpretability, and policy adherence. Perspective-aware AI approaches this by transforming a user’s aggregated personal data into a structured identity model called a \emphChronicle: a temporal knowledge graph that represents and grows with the user. Chronicles support the secure disclosure of context across federated networks. A Chronicle holder may expose a queryable, authorized view that a third-party agent may consult without centralizing anyone’s data. This paper explores the problem of minimum-necessary disclosure across domain boundaries: when a requester’s agent queries a Chronicle, how can the system constrain its response to release only what the requester’s relationship, stated purpose, and specific task require? We propose \textbfProvenance Preserving Chronicles (PPC), a federated protocol that compiles each holder’s Chronicle into a compact \emphauthorized evidence subgraph governed by one rule: \emphshare no more than the request requires. Holders keep local sovereignty; an access controller projects relationship-aware views over domain-expert ontologies; and a two-phase flow returns provenance-linked text first, releasing raw artifacts only after explicit holder approval. We frame the problem, map gaps in blockchain, P2P, and holder-sovereign designs, define the core constructs, and sketch the protocol with an explicit threat model.

[AI-140] Building AI That Works: ESnets Prag matic Approach to AI-Driven Operational Excellence

链接: https://arxiv.org/abs/2607.22948
作者: Bin Dong,Sukhada Gholba,Brooklin Gore,Shawn Kwang,David Mitchell,Samuel Oehlert,Garrett Stewart,Brendan White,Luke Baker,Ed Balas,Britt Gathright,Chin Guok,Jon-Paul Heron,John MacAuley,Scott Richmond,Chris Robb,Chris Tracy,Kesheng Wu
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and to address persistent operational pain points in the Network Operations Center (NOC) workflow. ESnet operators experience slow retrieval from siloed data sources, incidents described in lengthy and difficult-to-parse tickets, and context loss across shift handoffs. These challenges increase cognitive load and prolong incident resolution times. ORBIT therefore targets routine automation, cross-source synthesis, and actionable insights delivered directly within operators’ existing tooling. ORBIT is an agentic AI system integrated into ServiceNow, ESnet’s primary incident management platform. The design uses a modular, layered architecture comprising a centralized reasoning hub, tool access via MCPs for ESnet data sources, a semantic search layer, and an operator-facing chat interface. To manage the complexity and stochasticity of the AI toolchain, ORBIT follows industry best practices by structuring task logic as versioned, tested “skills” that guide the system in performing bounded responsibilities. This improves reliability and predictability compared to fully unconstrained agent behavior. Key results show that ORBIT successfully delivered all six initial tasks, and the architecture enabled rapid development of two additional tasks proposed by NOC engineers. We observed strong organic adoption of general-purpose infrastructure components, especially the chat interface and LiteLLM model gateway, including high request volumes from outside the project. Experiments with skills indicate that this approach can reduce task completion steps while eliminating observed error modes. Subjects: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22948 [cs.NI] (or arXiv:2607.22948v1 [cs.NI] for this version) https://doi.org/10.48550/arXiv.2607.22948 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-141] Invariant Discovery for Networked Systems

链接: https://arxiv.org/abs/2607.22944
作者: Hongyu Hè,Alexander Krentsel,Sylvia Ratnasamy,Maria Apostolaki
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Symbolic Computation (cs.SC)
备注: 8 pages, 4 figures, 1 table

点击查看摘要

Abstract:Invariants, the relations expected to hold among measured signals of a network, underpin applications from verification to traffic generation, telemetry imputation, and input validation, yet writing them by hand demands rare expertise in both formal logic and networking. Automatic miners can help but fall short on two fronts: they still require the hardest input (the grammar of admissible invariants) and they learn only exact, hard'' rules, struggling with real-world approximation caused by inherent noise in data. LLMs are tools that can provide semantic reasoning over data, but are non-deterministic and opaque in their learning. Our key idea is to partition the invariant search problem into an AI-driven grammar discovery’’ problem, followed by a statistics-driven ``search’’ problem within the learned grammar. Taken together, this allows non-deterministic, hallucination-prone AI to help produce auditable invariants with formal guarantees. We design and implement such a system, Autogram, and evaluate it on both public and production telemetry data, recovering expert-derived invariants with high coverage and low false positives. We close with discussion on open problems on the path toward fully open-ended discovery.

[AI-142] Design Theater: A Benchmark for Generative UI AAAI

链接: https://arxiv.org/abs/2607.22928
作者: Kashif Imteyaz,Kaif Imteyaz,Nakul Rajpal,Kaif Shaikh,Michael Muller,Saiph Savage
类目: Artificial Intelligence (cs.AI)
备注: accepted at AAAI/AIES

点击查看摘要

Abstract:Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces. Alongside the interface, these tools generate user-facing design rationales that explain their layout, accessibility, and design choices. However, it remains unclear whether these stated rationales are actually reflected in the interfaces they produce. We call this disconnect ``Design Theater’': plausible and confident design rationales that have little relationship to the actual implementation. To study this phenomenon, we introduce a benchmark and three metrics for measuring Design Theater. The benchmark includes 24 UI generation tasks spanning structural, styling, and functional design requirements. Using this benchmark, we evaluate 120 interfaces created by five generative UI tools. On average, over 25% of user-facing design rationales are not implemented in the generated interface, and the implementation failure increases to 34% for functional requirements. Tools recognize roughly half of the UX principles embedded in prompts (mean = 0.54), with four of five tools implementing 6% or fewer functional principles. We also measure interface similarity across tools and find convergence in visual appearance and layout organization, with greater variation in color choices. Overall, we contribute: 1) the concept of Design Theater; 2) a benchmark with metrics for assessing whether the stated reasoning of generative UI tools is reflected in their implementations; 3) and findings from a systematic evaluation of these tools. We discuss what these findings mean for the design and evaluation of generative UI tools.

[AI-143] SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

链接: https://arxiv.org/abs/2607.22926
作者: Mahdi Eslamimehr
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 25 pages, 5 figures, 5 tables

点击查看摘要

Abstract:High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authorization-separated architecture in which credible catastrophic-enablement risk constrains admissibility before utility, latency, or commercial objectives are considered. It combines signed release manifests, diverse detectors, robust risk envelopes, least-risk defaults, output checking, three-valued monitoring, protected audit chains, containment, and rollback. Formal results establish safety priority, conservative detector bounds, monotone release gating, tamper-evident records, and an authorization cut; two PRISM abstractions verify authorization separation and lifecycle invariants under explicit assumptions. A frozen, vendor-symmetric study sent 84 cases to each of four GPT, four Claude, and two Gemini snapshots: 840 calls yielded 794 target responses, 46 provider errors, and 449 successful judgments covering 375 responses. Eight snapshots had complete judged domain coverage. Harmful-compliance estimates were low; variation arose mainly from benign utility and safe redirection. Seven multiplicity-adjusted contrasts involving Claude, Gemini, or GPT-5 snapshots and the GPT-5 mini and GPT-5 nano snapshots were supported, while no tested contrast between the Claude or Gemini snapshots and GPT-5 or GPT-5.5 survived correction. The observed harmful-compliance range is a conservative, protocol-bound view from one generation per prompt with no tools, retrieval, history, or human adjudication; it is not an upper bound on operational assistance. A preregistered extension specifies how to test a wider best-worst gap using a locked split, repeated sampling, multi-turn and sandboxed-tool conditions, and domain-expert scoring. Comments: 25 pages, 5 figures, 5 tables Subjects: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) Cite as: arXiv:2607.22926 [cs.AI] (or arXiv:2607.22926v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.22926 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-144] How Well Can AI Generate Backlogs from App Mockups?

链接: https://arxiv.org/abs/2607.22902
作者: Andrea Lezcano Airaldi,Lourdes Romera,Walid Maalej
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: Accepted at the AIRE Workshop, IEEE 34th International Requirements Engineering Conference (RE) 2026

点击查看摘要

Abstract:Creating sprint backlogs requires considerable effort, as items such as epics, user stories, and tasks can be missed or inconsistently specified. We propose a multimodal approach to support backlog generation from visual app mockups, an artifact available at early project stages. We evaluate three prompting strategies on GPT-4o: a zero-shot baseline, Compositional Chain-of-Thought (CCoT) for vision-language reasoning, and a persona-driven prompt. We study seven app development projects across two countries and interview developers about the results. Overall, we observed that the baseline prompt favours recall over precision, whereas CCoT is more balanced, achieving average F1 scores of 52-66% for epics and user stories. Tasks were more challenging to generate accurately. Precision gains were most consistent when adding architectural context, particularly for backend tasks (precision gains up to 35%). Interviews with developers revealed that up to 26% of false positives were still considered useful, reflecting the creative and open-ended nature of backlog creation. To capture this, we propose a new measure called Revised Recall, which complements ground-truth evaluation with developer assessments. Our findings suggest that hybrid prompting with architectural context can assist backlog generation from early mockups, though results vary by item type and developer oversight remains necessary.

[AI-145] Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM -Generated Unit Tests ISSTA ISSTA113 ISSTA2026

链接: https://arxiv.org/abs/2607.22883
作者: Junda Zhao,Shurui Zhou,Eldan Cohen
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 24 pages, 7 figures, 12 tables. Accepted at ISSTA 2026; to appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA113

点击查看摘要

Abstract:While Large Language Models (LLMs) show great promise for automating unit test generation, recent studies suggest that the quality of generated tests can be negatively impacted when models are prompted with buggy code. This paper presents a new metric to quantitatively measure the “misguidance effect,” a phenomenon where buggy code steers LLMs toward generating tests that validate its erroneous behavior rather than expose it. Our analysis reveals that prompting LLMs with buggy code has a severe, twofold impact: it significantly increases “misguided tests” that assert incorrect behavior while simultaneously suppressing the generation of effective, bug-finding tests. We further corroborate this effect from a model-internal perspective, showing that buggy code skews LLMs’ preference toward tests that assert the same erroneous behavior. To counter this, we introduce and validate a specification-based unit test generation paradigm that replaces the code under test in the prompt with an LLM-generated specification docstring. Our results show that this paradigm effectively reduces misguided tests while substantially increasing effective tests, improves multi-round, feedback-driven test generation pipelines, and remains applicable to both buggy and bug-free code. Overall, these results suggest that specification-based prompting is a promising strategy for mitigating misguidance from buggy code in LLM-generated unit tests.

[AI-146] Do Coverag e and Mutation Scores of LLM -Generated Test Suites Correlate with Their Effectiveness? (Replicability Study) ISSTA ISSTA002 ISSTA2026

链接: https://arxiv.org/abs/2607.22880
作者: Junda Zhao,Shurui Zhou,Eldan Cohen
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 24 pages, 4 figures, 14 tables. Accepted at ISSTA 2026; to appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA002

点击查看摘要

Abstract:Recent advances in large language models (LLMs) have driven growing interest in using LLMs to automate test generation. Prior work commonly evaluates generated test suites using proxy metrics such as code coverage and mutation score. However, studies by Inozemtseva et al. and Papadakis et al. show that, for human-written tests, correlations among coverage, mutation, and real-bug detection can largely vanish once test suite size is controlled, raising concerns about the validity of evaluations based on proxy metrics. It also remains unclear whether these conclusions carry over to LLM-generated tests, given that prevailing LLM-based test-generation workflows differ substantially from traditional approaches. In this paper, we conduct a large-scale replication study of these two prior works using a wide range of test suites generated by a diverse set of LLMs, and re-examine the relationships among coverage, mutation, and real-bug detection effectiveness. Our findings diverge substantially from prior results. We show that the usefulness of coverage and mutation is highly context-dependent: in regression-style settings where the code provided to the LLM can be reasonably assumed bug-free, these metrics can provide meaningful signals when comparing across models; in another common scenario where the code-under-test may already be buggy and the goal is to expose the bug within the code-under-test, they no longer serve as reliable indicators. We also find little evidence that test suite size is a dominant confounder for correlations among coverage, mutation, and real-bug detection for LLM-generated tests. Based on these findings, we discuss how to interpret results from prior studies and provide actionable guidance for evaluating LLM-based test generation. Comments: 24 pages, 4 figures, 14 tables. Accepted at ISSTA 2026; to appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA002 Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) ACMclasses: D.2.5 Cite as: arXiv:2607.22880 [cs.SE] (or arXiv:2607.22880v1 [cs.SE] for this version) https://doi.org/10.48550/arXiv.2607.22880 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Related DOI: https://doi.org/10.1145/3832093 https://doi.org/10.1145/3832093 https://doi.org/10.1145/3832093 https://doi.org/10.1145/3832093 https://doi.org/10.1145/3832093 https://doi.org/10.1145/3832093 https://doi.org/10.1145/3832093 https://doi.org/10.1145/3832093 https://doi.org/10.1145/3832093 Focus to learn more DOI(s) linking to related resources

[AI-147] Spatial Prediction of Soil Microplastics and Organic Matter Using Graph Attention Networks

链接: https://arxiv.org/abs/2607.22875
作者: Anik Dev Nath,Md Al Amin,Bikash Kumar Paul
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 11 Pages, 8 Figures

点击查看摘要

Abstract:Accurate estimation of soil microplastics and organic matter is essential to assess ecosystem health and support sustainable land use. This study presents a graph-based deep learning approach using Graph Attention Networks (GATs) to model spatial dependencies among 91 georeferenced soil samples. By incorporating spatial coordinates, soil properties, and land use data, a two-layer GAT architecture was developed to capture local interactions. The final model showed strong performance, achieving RMSEs of 625.06 ( R^2 = 0.87 ) for microplastics and 0.43 ( R^2 = 0.91 ) for organic matter. However, cross-validation results revealed limited generalization, probably due to the small sample size and sparse graph structure. These findings demonstrate the potential of GATs for spatial soil prediction and underscore the need for dense datasets and improved graph connectivity.

[AI-148] Multi-primitive in-memory computing for Monte Carlo tree search

链接: https://arxiv.org/abs/2607.22869
作者: Tergel Molom-Ochir,Benjamin F. Morris III,Yintao He,Archit Gajjar,Giacomo Pedretti,Hai Helen Li,Yiran Chen,Jim Ignowski,Aishwarya Natarajan
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)
备注: 23 pages main text, 5 figures; 30 pages supplementary information

点击查看摘要

Abstract:Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge deployment. In-memory computing (IMC) is energy-efficient on regular workloads but has been considered incompatible with irregular multi-phase algorithms. We introduce phase-to-primitive decomposition, which reformulates each algorithmic phase as a hardware-native IMC primitive. Applied to MCTS, selection, expansion, rollout and backpropagation map to content-addressable memory, combinational logic, a resistive random-access memory (RRAM) crossbar and static random-access memory, keeping search on chip. At 22 nm with fabricated RRAM-array parameters, IMC-MCTS consumes ~60 mW for 9x9 Go, achieving 96x energy efficiency over a central processing unit (CPU) and 65x-2,059x over an H100 graphics processing unit (GPU). It reaches a European Go Federation rating within sample-size uncertainty of open-source Go engines (Pachi-UCT and Michi-C). The same substrate runs eight applications across four AI domains.

[AI-149] What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

链接: https://arxiv.org/abs/2607.22868
作者: Shawn Ray
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
备注: 26 pages, 8 figures. Extended version with complete proofs and additional experiments

点击查看摘要

Abstract:Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior. We separate three questions. First, relative to fixed oracle predicates, a deterministic gate enforces exactly the nonempty safety policies whose good prefixes its register model recognizes; policy nontriviality is undecidable with two decrementable counters but in PSPACE for a separable monotone fragment. Second, under a fixed exogenous law, Neyman-Pearson gives the exact false-block/miss frontier and conformal calibration gives a finite-sample marginal certificate, possibly via block-all. Third, once blocking changes future proposals, static scores and ungated trajectories need not identify the closed-loop frontier; a specified finite controlled model instead yields an occupancy program. Bounded representation attacks add a robustness margin, so benign calibration alone does not transfer. Experiments target these distinctions through static diagnostics, controlled-model enumeration, representation rewrites, and paired closed-loop reruns.

[AI-150] Coordinated Networking for On-Device Agent -Augmented Real-Time Communication

链接: https://arxiv.org/abs/2607.22854
作者: Goodsol Lee,Juheon Yi,Jinglu Wang,Haowen Xu,Saewoong Bahk,Yan Lu
类目: Artificial Intelligence (cs.AI)
备注: Accepted to USENIX NSDI 2027

点击查看摘要

Abstract:AI agents are enabling a new paradigm of agent-augmented real-time communication (RTC), where humans focus on high-level collaboration, while agents autonomously retrieve, analyze, and generate information in real time to support their interactions. These apps enable new experiences across various domains: for example, when corporate employees co-author a legal document, their agents can discuss and draft on their behalf, sparing them the burden of manually reviewing each other’s work. As existing cloud-based agents suffer from privacy risks and unscalable server costs, on-device agent-augmented RTC offers a promising alternative. However, this on-device paradigm introduces a new networking challenge: contention between concurrent traffic flows generated by humans (for live video streaming) and agents (for sending context files for analysis). We design HFS, a framework to ensure both high live video quality and low agent response latency in agent-augmented RTC apps. We achieve the goal through an app-guided multi-flow transport approach, where a unified app-layer orchestrator jointly controls the sending rates of live video and agent context flows based on their heterogeneous app requirements. Our prototype built atop WebRTC and this http URL demonstrates that HAFS outperforms baselines, achieving 1.5x higher video quality while reducing agent response time by 31%.

[AI-151] Disentangling Multi-View Scanning in Mamba for Network Traffic Anomaly Detection KDD2026

链接: https://arxiv.org/abs/2607.22829
作者: Xinglin Lian,Chengtai Cao,Ting Zhong,Fan Zhou
类目: Artificial Intelligence (cs.AI)
备注: Accepted by KDD 2026

点击查看摘要

Abstract:Network Traffic Anomaly Detection (NTAD) is a critical task in cybersecurity, yet timely and accurate anomaly detection remains challenging. Mamba has emerged as a particularly promising backbone for NTAD due to its linear-time complexity for long-sequence modeling. It further incorporates a dedicated multi-view scanning mechanism to enhance detection precision through complementary contextual cues. However, we identify a previously overlooked structural deficiency in multi-view Mamba scanning for NTAD: redundancy accumulation. Specifically, distinct scanning branches capture substantial view-invariant information, which is repeatedly amplified during multi-view fusion; conversely, view-specific information is diluted or even suppressed, leading to representation homogenization and multi-view degradation. To address this problem, we propose DisenMamba, a novel disentangled multi-view Mamba framework. DisenMamba reformulates multi-view scanning as a two-stage disentangle-then-fuse process that explicitly separates view-invariant and view-specific components prior to fusion. This design prevents the invariant information accumulation while preserving complementary multi-view cues, yielding more discriminative representations for subtle traffic anomalies. Extensive experiments demonstrate the effectiveness of DisenMamba, establishing a new disentangled multi-view Mamba paradigm. Code is available at this https URL.

[AI-152] From Hybrid Mechanistic–Data-Driven Modeling Toward Neuro-Symbolic AI: What Why and How

链接: https://arxiv.org/abs/2607.22811
作者: Moein E. Samadi,Andreas Schuppert
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO); Machine Learning (stat.ML)
备注:

点击查看摘要

Abstract:Hybrid mechanistic/data-driven models, which combine first-principles with learned components, are increasingly used in process engineering and scientific machine learning. Common hybrid modeling designs are specified primarily through their architectures and training losses, which offers a limited basis for a shared semantic interface to compare or verify them across domains, with comparatively little attention paid to epistemic uncertainty in the mechanistic part. We bridge hybrid modeling and neuro-symbolic (NeSy) AI by reconstructing these designs as instances of NeSy interface. The resulting translation, Hybrid-to-NeSy (H2N), places mechanistic knowledge on the language side, learned modules on the belief side, and validity domains together with constraints on the logic side. For each design, H2N then yields an explicit NeSy inference functional and a logic-belief decomposition. From this decomposition we derive two metrics: structural violation rate (SVR), measuring whether the learned belief respects the mechanistic structure; and belief dispersion (BD), measuring how concentrated the learned plausibility is, serving as a hybrid model’s epistemic uncertainty in its mechanistic part. We instantiate H2N on a case study of a structured hybrid model for binary classification under label noise and show that models with higher SVR and BD exhibit greater variability in held-out accuracy. Under structural distribution shift, H2N further quantifies a model’s uncertainty during extrapolations, whereas test accuracy reveals the same shift only post hoc. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO); Machine Learning (stat.ML) Cite as: arXiv:2607.22811 [cs.LG] (or arXiv:2607.22811v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.22811 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-153] OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge Intelligence

链接: https://arxiv.org/abs/2607.22805
作者: Keya Patel,Sajib Mistry,Sheik Mohammad Mostakim Fattah,Aneesh Krishna
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We propose OrchNAS, an energy-aware, personalised, federated edge intelligence framework that leverages a Neural Architecture Search Service to automatically design service-adaptive models for heterogeneous edge environments. The framework orchestrates the architecture search process on a server-side NAS service, enabling edge services to derive personalised architectures under device-level energy, computation, and memory constraints. We introduce an energy-aware global architecture search mechanism that learns a compact global representation across heterogeneous services. We develop an energy-efficient architecture selection mechanism that enables each service to derive a personalised subnet that satisfies its resource constraints via a progressive, greedy, energy-aware pruning strategy. We propose an energy-efficient personalised model optimisation scheme that updates service-adaptive parameters while preserving global representations, where a primal-dual optimisation mechanism enforces strict energy budgets during architecture adaptation. Experiments on real-world and benchmark datasets demonstrate the effectiveness of the proposed approach.

[AI-154] LithoFormer: A Robust Framework for Stratigraphic Inference via Transformers

链接: https://arxiv.org/abs/2607.22804
作者: Shwetha Salimath,Francesca Bugiotti,Sylvain Wlodarczyk,Sohaib Ouzineb
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 8 pages, 4 figures

点击查看摘要

Abstract:Accurate geological characterization of subsurface reservoirs from well log data is essential to support projects such as carbon capture and storage (CCS), geothermal development, and extraction of natural resources. Existing automated techniques for geological characterization primarily use sliding-window classification, which limits their ability to understand broader geological contexts, often leading to misaligned formation layers. To overcome these limitations, we introduce LithoFormer, a robust framework for stratigraphic inference using a Seq2Seq transformer model that ingests entire multivariate well logs in a single pass. The framework utilizes a channel-independent PatchTST backbone enhanced with rotary positional embeddings (RoPE) to capture long-range geological dependencies across entire multivariate well logs. A decoupled multi-task head is employed to jointly predict geological zonation and precise boundary probabilities, while a geology-informed loss function enforces physical constraints such as the Law of Superposition. Validated and deployed on three real-world datasets, LithoFormer demonstrates a 90% reduction in median boundary error and eliminates stratigraphic order violations compared to traditional sliding-window baselines. It also achieves a 80% reduction in manual expert labor and eliminates stratigraphic inconsistencies, providing a scalable and reliable solution for large-scale subsurface modeling.

[AI-155] Physically Verifiable Evidence and LLM -Based Reporting for Bearing Fault Diagnosis

链接: https://arxiv.org/abs/2607.22797
作者: Yuntong Chen,Jianyu Liu,Guobin Zhao,Ziang Wang,Chao Chen,Ju Huang,Xitian Tian,Lijiang Huang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon. Current intelligent fault diagnosers fail this standard in two ways. Their standard output, a class label with a softmax confidence score, is an internal statistic of the classifier, offering nothing checkable against independent physical knowledge; and the growing use of generative language models in maintenance reporting adds a second risk: hallucinated content entering reports on which decisions rest. Taking bearing fault diagnosis as the testbed, this work addresses both problems from the output side. The proposed Diagnostic Evidence Network (DENet) is an encoder-agnostic multi-task framework extending the output to a structured evidence record: the classification, a predicted characteristic frequency comparable against the theoretical value determined by bearing geometry and shaft speed, and a temporal localization of transient impulses inspectable on the raw waveform. Across four encoders and three public datasets, this evidence incurs no statistically significant accuracy cost, with a frequency error of about 6 Hz on 1,024-point segments where spectral estimation is structurally inapplicable. Centrally, the deviation between predicted and theoretical frequency constitutes a label-free, inference-time validation signal: it detects misclassifications with AUROC values of 0.970 and 0.871, and remains discriminative in the high-confidence regime where confidence-derived detectors are blind. Finally, a QLoRA-adapted language model is constrained to translate, but never generate, diagnostic content, reducing unsupported-claim rates from 10-12% to 2% and eliminating fabricated quantities.

[AI-156] FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow

链接: https://arxiv.org/abs/2607.22788
作者: Zhilin Huang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注: 11 pages, 1 figures

点击查看摘要

Abstract:AC optimal power flow determines the minimum-cost generation dispatch under nonlinear power balance constraints and is solved thousands of times daily in electricity market operations. Learning a direct mapping from load conditions to OPF solutions can accelerate this computation, yet with deepening renewable penetration, a single optimal dispatch is no longer sufficient. Operators require a characterization of the distribution of feasible near-optimal solutions for risk quantification, sensitivity analysis, and multi-objective trade-off assessment. Supervised neural networks provide fast point predictions but cannot capture this conditional distribution. Diffusion-based generative models can sample diverse solutions in principle, yet existing methods operating in the raw state space exhibit degraded solution quality and fail to scale beyond medium-sized systems. We identify the root cause as the conflation of two distinct tasks within a single model. Compressing the high-dimensional OPF solution manifold is one task, and learning the conditional mapping from loads to that manifold is another. This paper presents FMOPF, a framework that resolves this conflation by decoupling compression from generation through latent flow matching and by explicitly modeling load-state coupling through a Constraint-Aware Interaction Prior Network. Experiments on four IEEE test systems demonstrate that FMOPF provides the most effective Newton-Raphson warm starts, achieves the lowest tail risk among generative methods, and is the first such method to scale to systems with several hundred buses while preserving full feasibility. Ablation studies confirm that the latent generation pipeline is a necessary condition for physical feasibility and that the interaction prior functions as a late-stage tail-risk controller.

[AI-157] Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs

链接: https://arxiv.org/abs/2607.22786
作者: Ilia Sobakinskikh,Paul Alexander Bilokon
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC); Performance (cs.PF); Computation (stat.CO)
备注:

点击查看摘要

Abstract:In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series. The financial time series are price series such as asset prices. Unfortunately, the data is often with errors or outliers that make the downstream data processing tasks useless, unstable or even harmful. Moreover, the amount of financial time-series data has been significantly increasing. Hence, there is a need for better data-cleaning methods in terms of accuracy and in terms of processing speed. Transformers as a neural network architecture have achieved superior performances in many tasks such as Natural Language Processing and Computer Vision. Time series modelling and especially anomaly detection tasks can benefit from the features of transformers architecture in multiple ways, including the capacity to capture long-range dependencies and interactions. Increasingly powerful hardware, such as field-programmable gate arrays (FPGAs), have seen increasing usage in recent years due to their reconfigurability and high performance. They can be efficiently utilized to speed up the computations of the Transformer architecture. We explore different Transformer architectures for time series modelling and how they can be efficiently implemented on an FPGA board (PYNQ-Z2). In particular, we examine the application of Transformers to detect anomalies in time series and we show how they can be efficiently implemented on an FPGA board to minimize latency. The code is available at this https URL Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC); Performance (cs.PF); Computation (stat.CO) MSC classes: 68W35 (Primary), 68T07, 62M45 (Secondary) ACMclasses: I.2.6; C.3; I.5.4 Cite as: arXiv:2607.22786 [cs.LG] (or arXiv:2607.22786v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.22786 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: The Journal of FinTech, Vol. 5, No. 1, 2550001 (2025) Related DOI: https://doi.org/10.1142/S2705109925500014 Focus to learn more DOI(s) linking to related resources Submission history From: Paul Bilokon [view email] [v1] Fri, 24 Jul 2026 12:09:19 UTC (1,455 KB) Full-text links: Access Paper: View a PDF of the paper titled Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs, by Ilia Sobakinskikh and Paul Alexander BilokonView PDFHTML (experimental)TeX Source view license Current browse context: cs.LG prev | next new | recent | 2026-07 Change to browse by: cs cs.AI cs.AR cs.DC cs.PF stat stat.CO References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) IArxiv recommender toggle IArxiv Recommender (What is IArxiv?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[AI-158] What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation

链接: https://arxiv.org/abs/2607.22781
作者: Minwoo Yu,Young-guk Ha
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temporal graph models, for example, can achieve high future-link AUC while basic graph statistics remain difficult to recover from the same representation. We identify one source of this gap in the weighted averaging used by standard attention: when an evidence pattern is repeated, the numerator and denominator grow at the same rate, so inputs with different amounts of accumulated evidence can produce the same aggregate. We propose Mass-Aware Attention (MAA), which generalizes standard L1 normalization to an Lp family. Under repetition, MAA makes the numerator and denominator scale at different rates, retaining the effective number of contributing inputs in the representation magnitude. It adds no supervision, parameters, hidden dimensions, or explicit count features, and recovers standard attention at p=1. Across four continuous-time dynamic graph models and three datasets, MAA improves future-link AUC in 11 of 12 model-dataset cells. Linear recovery from the same hidden representation increases by 4.49% on average, and preferential-attachment recovery improves in all 12 cells after family-wise correction. We also observe consistent evidence in marked temporal point processes, temporal knowledge graphs, retrieval-augmented generation, and spatio-temporal point processes. Information accessibility and task utility remain distinct: NLL improves in MTPP, ranking is largely preserved in TKG, additional information in RAG does not improve the diagnostic head, and downstream LayerNorm can erase the signal in STPP. These results position MAA as a general normalization principle for improving predictor-facing representation informativeness by controlling repetition invariance in standard attention. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22781 [cs.LG] (or arXiv:2607.22781v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.22781 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-159] Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control

链接: https://arxiv.org/abs/2607.22779
作者: Federico Del Pup,Elisa Tentori,Manfredo Atzori
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: GitHub repository: see this https URL

点击查看摘要

Abstract:Hand gesture recognition via surface electromyography (sEMG) is fundamental to prosthetic control. In this field, deep learning approaches have become the gold standard. However, current architectures struggle to scale; model performance typically decreases as the number of hand movements increases. Performance degradation is tied to the increased statistical complexity of decoding expanded gesture sets and compounded by the limitations of state-of-the-art methods, which primarily rely on low-latency unimodal convolutional architectures. Convolutions operate locally, limiting model’s ability to capture long-range sequential patterns. Unimodal setups cannot leverage complementary information from coordinated signals characterizing movement execution, such as inertial and eye-tracking data. These limitations motivate architectures that integrate local and global features across multimodal physiological sequences. To bridge this gap, this study introduces EMG-CrossFormer, an end-to-end hybrid convolutional-transformer for seamless multimodal integration. EMG-CrossFormer combines representations from an arbitrary number of unimodal encoders through cascaded cross-attention fusion layers, and decodes the fused representations using learnable gesture queries. EMG-CrossFormer was evaluated on four NinaPro datasets (DB2, DB3, DB7, and DB10) and benchmarked against six state-of-the-art models using an increasing number of modalities. Using only sEMG, EMG-CrossFormer achieved mean accuracies of 72.33%, 52.48%, 79.16%, and 73.49% on DB2, DB3, DB7, and DB10, respectively. Incorporating inertial signals improved performance to 90.66%, 80.40%, 92.79%, and 92.06%. These results show that joint local-global feature modeling improves sEMG-only decoding and that multimodal fusion substantially amplifies this benefit, underscoring the value of both design principles for complex hand gesture recognition.

[AI-160] LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning

链接: https://arxiv.org/abs/2607.22777
作者: Chen Wang,Boming Kang,Qinghua Cui
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid sequences, their training objectives do not explicitly constrain the model to learn three-dimensional residue contacts formed after folding . Here, we introduce LC-SEPLM (Long-range Contact-supervised ESM Protein Language Model), which adapts ESM2 with LoRA and long-range residue-pair contact supervision while retaining sequence-only downstream inference. Pair-specific queries use cross-attention over the complete sequence to extract global sequence context associated with long-range spatial contacts. To expose the model to diverse structural information, we trained LC-SEPLM on 500,000 AlphaFold Swiss-Prot proteins. In downstream evaluation, LC-SEPLM improved all eight protein-level tasks relative to ESM2. The largest gain occurred in remote-homology recognition, where macro-F1 increased from 0.6122 to 0.6769 (+0.0647, or 6.47 percentage points). On the official ESM-S EC benchmark, LC-SEPLM also outperformed ESM-S with a maximum absolute gain of 0.1771. These results support residue-pair contact supervision as a bounded route for introducing structural information into protein sequence representations while preserving sequence-only inference.

[AI-161] DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning

链接: https://arxiv.org/abs/2607.22769
作者: He Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data. Existing dynamic data scheduling methods face critical limitations in industrial-scale pretraining and supervised fine-tuning (SFT): data selection incurs prohibitive O(N) costs on terabyte-scale corpora, mixture optimization schemes introduce severe I/O bottlenecks or require training auxiliary reference models, and sample-level reweighting strategies rely on loss signals that conflate noise, difficulty, and novelty. We present DomainPilot, a domain-level loss-guided two-stage data mixture optimization framework. DomainPilot introduces token-level domain loss monitoring to capture per-domain learning dynamics during training without halting the data pipeline. Building on these signals, we propose a Scaling Law guided coarse optimization stage that fits domain-specific convergence curves and derives a principled prior for mixture adjustment. A subsequent Mixing Law guided fine optimization stage refines the mixture by modeling cross-domain interaction effects through controlled sweep experiments. The entire mechanism is realized via a patch-based architecture that injects domain-aware loss computation into existing training frameworks (e.g., MindSpeed/Megatron-LM) with only ~30 lines of framework-specific adapter code. We validate DomainPilot on the Qwen3-1.7B model during SFT. Compared to the original data mixture, our optimized mixture achieves improvements of +2% on MMLU-Redux, +1.8% on AIME24, +3.8% on LiveCodeBench v5, and +3.6% on BFCL v3, without increasing total data volume or training cost. These results demonstrate that domain-level training signals provide an effective, lightweight alternative to expensive data selection or auxiliary model training for mixture optimization. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22769 [cs.LG] (or arXiv:2607.22769v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.22769 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-162] Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

链接: https://arxiv.org/abs/2607.22766
作者: Yunting Song,Matthew Watson,Peter Grabowski,Jun Qin
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks, and systemic human annotation errors. Standard dataset auditing methods, such as semantic deduplication or LLM-as-a-judge, struggle to capture the actual predictive impact of individual records and often miss deep functional rule clashes. To address this, we introduce a scalable, inference-only data valuation pipeline that approximates the Shapley value without iterative model retraining. By mapping semantic k-NN neighborhoods into a directed graph, our framework evaluates data utility directly through a reference LLM’s probability distribution using zero-shot and one-shot conditional log-likelihood shifts. Our pipeline then translates these predictive influence scores into localized advantage metrics to isolate gradient-conflicting records. We demonstrate the pipeline’s efficacy in sanitizing two heavily vetted alignment datasets. First, applying our pipeline to the HelpSteer2 dataset reduced the manual audit search space by 99.1%, successfully uncovering falsely-labeled records across diverse failure modes. Second, applying our automated audit strategy to Anthropic’s HH-RLHF training and evaluation splits identified thousands of hidden safety and factual preference inversions. Crucially, by extending this audit to the evaluation split, we expose severe vulnerabilities in current benchmark integrity: highly capable models frequently predict the safer or more helpful response, only to be penalized by objectively flawed human ground-truth labels. Overall, our work provides a mathematically grounded, highly efficient diagnostic tool to uncover human label failures, sanitize evaluation benchmarks, and ensure the integrity of LLM alignment data.

[AI-163] Hierarchical Grading in Large Language Models

链接: https://arxiv.org/abs/2607.22757
作者: T. Shaska
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training objective. The construction extends the theory of graded neural networks and graded transformers to autoregressive language models while preserving expressive power, asymptotic computational complexity, and inference cost. The governing geometric picture is that of geometric invariant theory. The benefit of a grading is expressed by a Kempf–Ness functional on the grading torus; the grades that improve upon the uniform architecture form an open convex cone whose membership is decided by a Hilbert–Mumford-type criterion pairing a grade direction against two measurable profiles of the target and the data; the optimal grades are the coincidence point of two moment maps, given in closed form; and the ordinary transformer appears as a semistable isotropic point on the boundary of the cone: one member of a larger graded family rather than a distinguished optimum. Separately, for level-stratified targets we prove a minimax separation between the graded prior and its absence: over all estimators the risks of the graded and uniform target classes separate throughout an explicit window of sample sizes, by a factor that decays exponentially in the number of levels under geometric stratification. Both profiles are estimable offline, so the optimal grades solve a convex program certified before training begins. Because the grading is absorbed into the learned parameters after training, every GLLM compiles to a standard transformer of identical architecture and inference complexity. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) ACMclasses: I.2.7; I.2.6; F.2.2; H.1.1 Cite as: arXiv:2607.22757 [cs.LG] (or arXiv:2607.22757v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.22757 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-164] Commitment To Cooperation With Self-Negotiated Contracts

链接: https://arxiv.org/abs/2607.22750
作者: Tim Wyse,Kaitlin Bustos,Yulia Volkova,Max Kleiman-Weiner
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:As AI agents operate with increasing autonomy in a multi-agent world, they will need to learn to cooperate with other agents and with humans to generate mutual benefits. However, cooperation is a challenge because the costs of cooperation are often incurred early on, but the benefits are only realized later, creating an incentive to defect. How can AI agents cooperate with commitment? Here, we draw on inspiration from legal institutions and contracting that human societies have used to solve principal-agent problems of this kind. Contracts provide observable representations of agreements that enable credible commitments through the enforcement of terms. We study the role of contract-based cooperation using LLM-based agents in \CT, a spatial-temporal game that combines bargaining with navigation towards a goal. We study a suite of contract representations that range from formal contracts that compile to code to natural contracts that require reinterpretation. We evaluate agents with a range of LLM backbones using different sizes and providers. We find that self-negotiated contracts can improve cooperative outcomes beyond what is possible with regular trading.

[AI-165] Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence

链接: https://arxiv.org/abs/2607.22748
作者: Zhaowen Fan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE)
备注: 22 pages, 6 figures, 2 tables

点击查看摘要

Abstract:Modern neural networks primarily adapt through parameter modification within predefined computational structures. While recent methods introduce modularity, conditional computation, and parameter-efficient adaptation, they generally do not distinguish computational capability from computational accessibility as separate adaptive variables. This work introduces Accessibility Plasticity, a principle of adaptive computation in which systems adapt not only by changing what computation exists, but also by reorganizing which existing computations can interact and participate. We formalize Accessibility Plasticity through a relationship-based operational realization and establish a reuse-first hierarchy of adaptation, where accessibility modification precedes more costly capability and structural changes. A proof-of-concept evaluation on sequential learning tasks shows that accessibility adaptation can reduce capability modification while maintaining comparable task performance. These results suggest accessibility as a distinct adaptive dimension and provide a foundation for future dynamic neural systems whose computational relationships evolve with changing environments.

[AI-166] Spatial Reasoning in LLM Game Agents : Impact of Causal Context and Multi-Step Planning

链接: https://arxiv.org/abs/2607.22732
作者: Mohit Jiwatode,Ronja Fuchs,Robin Schmöcker,Bodo Rosenhahn,Alexander Dockhorn
类目: Artificial Intelligence (cs.AI)
备注: To be published at COG 2026

点击查看摘要

Abstract:LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and evaluates whether causal prompt augmentation and multi-step planning can improve win-rates while managing response latency. Using the open-source Qwen3 model family, we conduct experiments across varying model scales, reasoning modes, and planning horizons. We further introduce a focused GVGAI benchmark consisting of three custom games with five difficulty levels to isolate spatial navigation. The evaluation follows two paradigms: an initial ``positioning experiment’’ to test an agent’s ability to find its exact coordinates, and a study of game-play success. Our results show that while larger models with an enabled thinking mode identify their positions more accurately, overall performance in coordinate matching remains limited for smaller models. Win rates decrease as game levels and layout complexity increase, validating the benchmark’s difficulty scaling. Integrating causal context into the prompts tends to improve the agents’ success rates, particularly for bigger models. While enabling thinking mode and longer planning horizons significantly improve performance, multi-step planning further reduces mean per-step response times, offering a practical trade-off between reasoning depth and execution speed.

[AI-167] Progress-conditioned Group Policy Optimization for Long-Horizon Agent ic Tasks

链接: https://arxiv.org/abs/2607.22724
作者: Kaibing Yang,Guangfeng Cai,Shengtian Yang,Shuo He,Yu Li,Mengyi Liu,Pengwei Chen,Jun Xu,Lei Feng
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison can suffer from a sampling imbalance: repeated or low-effect actions dominate the high-probability region of the policy while useful state-changing actions remain under-sampled. This imbalance produces many all-failed rollout groups, where outcome rewards provide no direction for correcting the policy. Together, these effects can form a self-reinforcing credit trap: failure-dominated sampling yields no outcome-based correction, allowing repeated low-effect actions to persist. To break this loop, we propose Progress-conditioned Group Policy Optimization (ProGPO), which uses first-visit observation coverage only when all samples in a group receive zero outcome reward. Specifically, within such groups, ProGPO assigns higher relative advantages to trajectories or steps that visit more new states since reaching new observations is a prerequisite for task success. Experiments on two challenging agentic benchmarks, ALFWorld and WebShop with Qwen2.5-1.5/7B-Instruct, show that ProGPO consistently improves over group-based baselines, with particularly large gains on hard tasks.

[AI-168] SEGRA: Structured Experience-Guided Graph Reasoning Agent for Gremlin Based Question Answering

链接: https://arxiv.org/abs/2607.22713
作者: Saiyue Lyu,Mariam Dundua,Vishaal Kapoor,Sarthak Ahuja,Neda Kordjazi,Evren Yortucboylu,Harsh Amin,Rebecca Steinert
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Enterprise IT support knowledge graphs capture rich relationships among cases, users, devices, symptoms, taxonomic categories, root causes, and historical resolutions. Yet querying them in Gremlin requires knowledge of graph schemas, traversal semantics, edge directionality, and property-graph-specific constraints, making them difficult for non-expert operators to use. We introduce SEGRA, an experience-guided agent for enterprise text-to-Gremlin question answering. SEGRA integrates intent routing, schema- and taxonomy-grounded query generation, multi-shot decomposition, execution-aware verification, and a curriculum-bootstrapped skill library that reuses verified query patterns. On an enterprise IT support benchmark, SEGRA achieves a 7.0\times higher mean judge score than backbone-only chain-of-thought prompting. Its skill library further reduces LLM calls by 20% and dollar cost by 18% relative to SEGRA without skills, while preserving answer quality. These results show that schema-grounded agent design and reusable execution experience improve both accuracy and efficiency for enterprise graph QA.

[AI-169] CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents

链接: https://arxiv.org/abs/2607.22711
作者: Mingwei Zheng,David OBrien,Siwei Cui,Pardis Pashakhanloo,Rajdeep Mukherjee,Myeongsoo Kim,Sachit Kuhar
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:

点击查看摘要

Abstract:LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. However, the conventional append-only trajectory architecture found in practice tightly couples file-read actions with their observations, capturing snapshots that become permanently fixed in the chronological history. As files change through agent edits or concurrent human modifications, these snapshots become stale, causing reasoning errors and causing agents to redundantly re-read files, with each re-read appending yet another copy to the trajectory. To mitigate this, we propose CORVUS, a novel trajectory architecture that decouples file-read actions from their observations by maintaining a synchronized registry of relevant files and injecting only their current contents at each reasoning cycle. This structural change produces significantly lighter-weight trajectories that remain synchronized with the actual codebase state by construction, eliminating redundant file copies and stale snapshots that bloat conventional trajectories. We evaluated CORVUS on SWE- POLYBENCH_VERIFIED and SWE-BENCH PRO across four LLMs, achieving 9-50% reduction in average input tokens per task, 15-32% shorter final prompts, and up to 37% fewer reasoning cycles while maintaining comparable pass rates.

[AI-170] RoleMix: Unifying Sequential and Non-Sequential Features via Semantic Tokenization for Post-Click Conversion Rate Prediction KDD

链接: https://arxiv.org/abs/2607.22700
作者: Wenan Wang,Qin Zhao,Zhixiang Lu
类目: Artificial Intelligence (cs.AI)
备注: KDD Cup 2026 Tencent UniRec Challenge

点击查看摘要

Abstract:Post-click conversion rate (PCVR) prediction is central to industrial recommendation, but remains challenged by the structural mismatch between sparse, unordered multi-field features and long, domain-specific behavior histories. Existing models often process these signals through separate pathways and fuse them late, weakening semantic roles and limiting cross-signal refinement. We propose RoleMix, a unified interaction architecture that represents sequential and non-sequential evidence through a shared, role-preserving token interface. Non-sequential fields are converted into explicit semantic tokens that preserve user, item, pairwise, dense, contextual, and cross-feature roles, while long behavior domains are compressed into item- and context-aware sequence-query tokens through two-stage hierarchical window attention. The resulting global, semantic, and sequence-query tokens are jointly refined by stacked UniMixing-Lite blocks for PCVR prediction. On the large-scale KDD Cup 2026 Tencent UniRec Challenge, RoleMix achieves 83.648% online AUC, outperforming the official industrial baseline by 1.953%. Ablation studies show that semantic tokenization yields the largest isolated gain, highlighting a key principle for large-scale PCVR modeling: preserving field semantics at the token-interface level is as important as scaling the interaction backbone.

[AI-171] Similarity All The Way Up: Multilingual Generalization in LLM s Relies on Language-Level Similarity Structures

链接: https://arxiv.org/abs/2607.22699
作者: Supantho Rakshit,Adele Goldberg,Henry Conklin
类目: Artificial Intelligence (cs.AI)
备注: 8 pages, 3 figures, submitted to CogSci 2026

点击查看摘要

Abstract:As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data. To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs’ representations capture the hierarchical similarity structure between distinct languages. Strikingly, we show LLMs’ latent representations largely recover the hierarchical structure of the Indo-European language family tree – grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.

[AI-172] PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window

链接: https://arxiv.org/abs/2607.22695
作者: Ryan Thornton,Mir Mehedi Ahsan Pritom,Maanak Gupta
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 9 pages, 3 figures, 3 tables. Submitted to IEEE CARS 2026

点击查看摘要

Abstract:Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment. While this automation provides clear utility, completing these tasks often requires the insertion of Personally Identifiable Information (PII), strings of information that uniquely identify some individual, raising privacy concerns. However, ethics has prevented the curation of a public, authentic dataset of PII. Without an appropriate dataset, it is difficult to quantify privacy risks. Thus, we introduce the PANOPTICON pipeline and dataset. The dataset, generated by Meta’s Llama-3.1-8B-Instruct model, contains 67, 718 prompts, intended for the models context window, containing PII spans derived from 9,674 publicly available synthetic user profiles. We measure lexical diversity and S-BERT diversity of the created dataset to evaluate realism. Finally, we present a case study showcasing the utility of PANOPTICON data for understanding Prompt Inversion Attacks (PIAs). PANOPTICON thus emerges as the first benchmark dataset for studying PIAs over private corpora, providing a foundation for future LLM privacy research.

[AI-173] Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models

链接: https://arxiv.org/abs/2607.22694
作者: Wenjie Fan,Bin Ma,Dong Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Attention collapse in autoregressive language models – manifested as repetitive token loops where the model becomes trapped in self-reinforcing attractors – is a persistent pathology that existing decoding-time heuristics fail to address at its root cause. We present a principled framework that penalises or compensates anomalous confidence arising from collapsed generation patterns, by comparing a token’s observed frequency against its corpus prior through an adjacent-conditional probability construction. The resulting self-normalising penalty ratio R=f(m,n,p)/f(np,n,p) requires no ad hoc standardisation and admits a closed-form logit offset with zero approximation error. The correction is isolated from the loss gradient and accumulated into a frozen output-layer bias via exponential moving average, enabling deployment as a repair mechanism for models that have already collapsed without requiring intrusive modifications to standard training pipelines. Experimental validation on a 1.5B-parameter model demonstrates that the frozen-bias mechanism can rescue a model already trapped in a collapsed attractor, reducing 2-gram repetition from 0.073 to near 0 while preserving generation quality.

[AI-174] Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

链接: https://arxiv.org/abs/2607.22692
作者: Anabela C. Areias,Catarina Botelho,António Farinhas,Areti Vassilopoulos,Dora Janela,Xin Tong,Nuno M. Guerreiro,Maya D’Eon,Fabíola Costa,Ricardo Rei
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existing safety approaches primarily detect risk but rarely shape how models respond as conversational risk unfolds. We developed a model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions. Synthetic conversations grounded in real-world mental health narratives were used to evaluate the architecture’s performance, tested with GPT-5-chat and Qwen3.5-27B, achieving high risk detection performance (specificity: 0.85 (95%CI: 0.78;0.91), sensitivity: 0.92 (95%CI: 0.88;0.95)) and increasing clinician-preferred escalation responses by 25.6–59.2pp while preserving rapport and connection. Performance remained stable across conversation length and generalized across both proprietary and open-source models. These findings demonstrate that clinically-grounded safety governance can extend beyond risk detection to improve how LLMs manage evolving mental health risk, providing a scalable framework for safer deployment across models.

[AI-175] HiLLTS: Zero-Shot Hierarchical LLM -Guided Traffic Signal Control for Sustainable Transportation

链接: https://arxiv.org/abs/2607.22691
作者: Yue Ding,Tendai Mukande,Mingming Liu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Urban traffic congestion significantly increases fuel consumption, greenhouse gas emissions, and commuter delays, resulting in substantial economic losses and environmental harm in modern cities. Traditional traffic signal control strategies such as fixed-time scheduling, actuated control, and reinforcement learning (RL)-based methods, offer different degrees of adaptability; however, RL-based methods can require extensive retraining, careful reward design, and substantial simulation data when transferred across networks or demand regimes. To address these challenges, we propose HiLLTS, an LLM-guided traffic signal control framework that employs a hierarchical three-layer architecture consisting of a central coordination agent, a district layer and multiple cluster-level intersection agents. Experimental results demonstrate consistent improvements in both congestion and environmental performance. Compared with the strongest non-LLM baseline in each scenario, HiLLTS reduces average waiting time by 36.73% under the low-congestion scenario and 14.71% under the high-congestion scenario, while reducing average CO2 emissions by 7.87% and 8.57%, respectively. Larger gains are observed against weaker baselines: under low congestion, HiLLTS achieves reductions of up to 18.00% in emissions and 62.07% in waiting time relative to Fixed-Time control; under high congestion, reductions of up to 28.89% in emissions and 40.36% in waiting time are observed relative to Max Pressure. The ablation study further validates the contribution of LLM-guided coordination over rule-based control

[AI-176] LazyMem: Retrieve Broadly Construct Selectively for Efficient Long-Term Agent Memory

链接: https://arxiv.org/abs/2607.22690
作者: Jing Yu,Yibo Zhao,Jiaming Zhang,Xiang Li
类目: Artificial Intelligence (cs.AI)
备注: 28 pages

点击查看摘要

Abstract:Long-term memory lets LLM agents reuse past interactions, but raw dialogue histories are verbose and information-sparse. Retrieving broadly improves evidence coverage yet overwhelms downstream reasoning with noise; compressing at write time reduces noise but irreversibly discards details the future query may need. We introduce LazyMem, which sidesteps this dilemma by deferring all memory construction to query time. A lightweight 4B model processes the retrieved candidate pool in overlapping parallel windows, selectively retaining and compressing only query-relevant content. The model is trained through supervised fine-tuning followed by group-based reinforcement learning with a format-gated composite reward that combines a rule-based action signal measuring selection accuracy with an LLM-judged quality signal measuring source faithfulness and query utility. On the LongMemEval benchmark, LazyMem-4B achieves an LLM-judge accuracy of 0.85 with only 213 memory tokens, 68.7 \times fewer than retrieval-only, and generalizes to LoCoMo (0.68) without target-domain training, while reducing mean latency over the prior query-time baseline. The 32B variant reaches 0.93, surpassing oracle-context references on aggregation-heavy question types. The code associated with this work is publicly available at this https URL.

[AI-177] Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents

链接: https://arxiv.org/abs/2607.22689
作者: Zedong Yu,Qianxing Li,Zhi Gao,Liuyu Xiang,Chenrui Shi,Yang Liu,Huiming Wu,Yujie Wei,Yuhao Fei,Yubo Fu,Zhaofeng He
类目: Artificial Intelligence (cs.AI)
备注: 15 pages, 5 figures. Project page: this https URL

点击查看摘要

Abstract:Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as clicking, typing, and scrolling on desktops and mobile devices. However, current agents scale poorly to long-horizon tasks: actions incur costly LMM inferences, and performance degrades as context grows. Humans divide such workloads among collaborators who complete sub-tasks in parallel. Yet parallel coordination among GUI agents has received little attention. To close this gap, we introduce ParaGUIBench, to our knowledge, the first benchmark dedicated to parallel execution and coordination of multiple GUI agents on separate desktop instances. It consists of three components: a multi-device Docker infrastructure with a shared file system; a dataset of 233 tasks spanning six task categories; and an evaluation system with efficiency metrics, including step reduction ratio and token cost. We further introduce ParaGUI, a planner-worker agent that decomposes GUI tasks and dispatches sub-tasks to concurrent workers on separate desktop instances. On ParaGUIBench, ParaGUI reaches a 46.4% success rate, outperforming the strongest serial baseline (Claude Sonnet 4.6) by 12.9 points while using roughly half the steps and less than half the tokens. These results show that parallel execution can improve both success rate and efficiency on decomposable, long-horizon GUI tasks, pointing to a direction worth further study.

[AI-178] DOSA: A Tree-Guided Self-Regressive Framework for Long Document Structure Analysis

链接: https://arxiv.org/abs/2607.22679
作者: Bohou Li,Benjamin Sowell,Mehul Shah,Mark Lindblad,Henry Lindeman
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In visually-rich documents, information is encoded not only in individual page objects such as tables, headers, and text blocks, but also in the structural relations among them, making document structure analysis fundamental to information retrieval and document understanding. However, accurately inferring such relations remains challenging in multi-page documents with long-range dependencies and heterogeneous layouts. To address this, we propose a tree-guided and self-regressive framework, termed DOcument Structure Analyzer (DOSA), for inferring relations among page objects and reconstructing document-level semantic trees. DOSA processes documents chunk-by-chunk, fusing visual, textual, and layout features for each page object and predicting hierarchical and ordering relations. The predicted relations are used to incrementally construct a semantic tree, which is then leveraged as structural context to guide inference on subsequent chunks. Experimental results on five benchmarks demonstrate the effectiveness of DOSA, with improvements of up to 4 F1 points and 19 TEDS points on DocHieNet, the most challenging multi-page hierarchy benchmark.

[AI-179] SetGo: Metadata Readiness for Scientific AI Datasets

链接: https://arxiv.org/abs/2607.22677
作者: Sean R. Wilkinson,Polina Shpilker,Wesley Brewer
类目: Digital Libraries (cs.DL); Artificial Intelligence (cs.AI)
备注: 6 pages; accepted at SSDBM 2026

点击查看摘要

Abstract:Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52-57% to 81-91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess-enrich-publish loop, with user involvement limited to supplying missing metadata values. Comments: 6 pages; accepted at SSDBM 2026 Subjects: Digital Libraries (cs.DL); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22677 [cs.DL] (or arXiv:2607.22677v1 [cs.DL] for this version) https://doi.org/10.48550/arXiv.2607.22677 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-180] AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

链接: https://arxiv.org/abs/2607.22671
作者: Rohan Naphade,Minzhou Pan,Bo Li
类目: Artificial Intelligence (cs.AI)
备注: 11 pages, 7 figures, 3 tables

点击查看摘要

Abstract:Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective. We present AIR-BENCH Live, a self-evolving successor to AIR-BENCH 2024. An automated update pipeline monitors government regulation and classifies new policies against the current four-tier risk taxonomy, either matching them to existing categories or proposing new granular categories. Then, a multi-agent, persona-driven prompt generation algorithm generates realistic, multilingual prompts with minimal human review, leaving room for improvement with modern jail breaking techniques. This algorithm is used to overhaul legacy prompts and generate prompts for new categories. In our current version, the pipeline has expanded the benchmark from 314 to 335 granular risks, with the 21 new categories drawing from 31 truly novel policy clauses across seven jurisdictions. Evaluating 14 recent models, we find a wide safety spread (from 0.17 to 0.89 among the models judged on their own behavior), that the modernized prompts are on average 0.06 points harder than the 2024 set, with the largest drops concentrated among the most compliant models, and that most models are modestly less safe on non-English prompts. By continuously absorbing new regulation and regenerating prompts, AIR-BENCH Live is designed to evolve alongside a fast-moving field.

[AI-181] Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance

链接: https://arxiv.org/abs/2607.22667
作者: Andrei Starodubov,Yaqub Aris Prabowo,Andreas Hadjipieris,Roberto Galeazzi,Ioannis Kyriakides
类目: Artificial Intelligence (cs.AI); Information Theory (cs.IT); Machine Learning (cs.LG); Robotics (cs.RO); Signal Processing (eess.SP); Systems and Control (eess.SY)
备注: 5 pages, 4 figures, submitted to the IEEE MetroSea 2026 Conference: Special Session 13: Object Detection, Tracking, and Sensor Fusion for Maritime Situational Awareness

点击查看摘要

Abstract:This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: instead of activating all sensors or repeatedly performing computationally expensive online expected-information-gain evaluation, a learned policy selects one tracking-relevant sensor at each decision epoch. A Bayesian sequential Monte Carlo tracker estimates the vessel state from noisy measurements and provides a belief representation for scheduling under nonlinear and non-Gaussian conditions. A Proximal Policy Optimization agent selects one of five sensors deployed in a georeferenced simulation of the CMMI Smart Marina testbed at Ayia Napa Marina, Cyprus. The agent observes belief-state, detection-history, coverage, sensor-geometry, and realized-information-gain features. The reward is defined as a realized-information-gain term gated by an observability mask. Final-test simulations compare the proposed framework with random single-sensor selection, always-on sensing using all sensors simultaneously, and the expected-information-gain sensor-selection baseline proposed in our previous work. Results show that the learned policy achieves tracking performance close to always-on sensing while activating only one sensor per decision time step and avoiding the computationally expensive online entropy search required by expected-information-gain selection.

[AI-182] Obliviate: Efficient Unlearning in Recommender Systems ICML2026

链接: https://arxiv.org/abs/2607.22665
作者: Tushar Prakash,Brijraj Singh,Niranjan Pedanekar,Narayan Chaturvedi
类目: Artificial Intelligence (cs.AI)
备注: Accepted at ICML 2026

点击查看摘要

Abstract:Machine unlearning is becoming increasingly critical in the context of data privacy regulations, particularly for recommendation systems that are directly trained on user interaction data. The goal of this work is to remove requested interaction data and their downstream influence from trained model while preserving recommendation quality, and to do so without incurring the substantial computational cost of full retraining. Existing approaches exhibit several limitations, including limited unlearning completeness and degradation in recommendation performance, while having substantial computational overhead. In this paper, we propose Obliviate, an efficient two-stage unlearning framework for recommender systems that achieves high unlearning completeness while maintaining good utility. In the first stage, we introduce a Low-Rank Unlearning Adapter (LUA), which employs a lightweight Hessian proxy to enable curvature-aware and efficient unlearning through localized low-rank adapters rather than full parameters. In the second stage, we propose Locality-Aware Calibration (LAC), a lightweight refinement stage that updates only the adapter parameters to improve the performance by enforcing unlearning via ranking-based objectives while preserving utility through knowledge distillation. Extensive empirical evaluations demonstrate that Obliviate achieves high level of forgetting with minimal loss in recommendation quality and at significantly reduced computational cost, offering a practical and scalable solution for large-scale recommender systems.

[AI-183] Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

链接: https://arxiv.org/abs/2607.22663
作者: Xingyu Mou,Zijin Huang,Tianze Zhang,Yuxin Ma,Lanning Wei,Zengfeng Huang,Da Zheng,Lun Du
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed-size blocks preserves parallel generation within each block while keeping the quadratic attention cost tractable. However, this efficiency comes with a structural limitation: tokens near the end of a block are generated without access to future cross-block context, and once a block is finalized, its uncertain predictions become irreversible context for all subsequent blocks. This creates a block boundary problem, in which uncertainty accumulates toward block boundaries and early mistakes propagate throughout later generation. To address this issue, we propose Multi-Block Editing (MBE), to mitigate this problem by editing decoded tokens based on cross-block context. Following this principle, MBE first proposes a training-free decoding algorithm to edit the decoded tokens in previous blocks, which is achieved by re-opening a full-attention window over selected blocks. Given the mismatched attention mechanism between block diffusion training and MBE inference, MBE further introduces a supervised Fine-tuning strategy, which equips the model with bidirectional attention masks that progressively expands the editing span. Furthermore, it also extends SGLang with a multi-shape CUDA Graph pool and fine-grained KV cache control to make these variable-length editing passes efficient in practice. Experiments on LLaDA2.1-Mini across 13 benchmarks show that training-free MBE outperforms all existing decoding baselines while maintaining comparable throughput, and MBE SFT further brings a performance gain of 2.7. The largest improvements appear on tasks requiring strong long-range consistency, including +13.3 on AIME 2025 and +5.9 on ZebraLogic, demonstrating the effectiveness of MBE.

[AI-184] CuraWeb: Joint Optimization of Quality Redundancy and Diversity for Web-Scale Pretraining Data

链接: https://arxiv.org/abs/2607.22662
作者: Peiguang Li,Yongwei Zhou,Juncheng Diao,Yuchun Fan,Jian Yang,Jianxiao Yang,Zhongda Su,Shuguang Jiao,Xiao Wei,Zhiye Zou,Gan Dong,Zhizhao Zeng,Rongxiang Weng,Jingang Wang,Xunliang Cai
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance. However, these pipelines typically rely on singular optimization objectives, which inevitably narrows distributional diversity and marginalizes long-tail knowledge, thereby restricting data coverage and underutilizing the vast potential of the open web. To address this limitation, we propose a novel curation paradigm that shifts from linear pruning to the joint optimization of quality, redundancy, and diversity. This framework synergizes dual-track cleaning (rule-based and model-driven) with hybrid deduplication (n-gram and semantic), while employing a multi-objective sampler to balance informational quality with distributional breadth. Applying this framework to Common Crawl, we construct CuraWeb, a 2T-token English corpus. Unlike existing resources, CuraWeb establishes an industrial-grade standard for data curation by recovering a more holistic data distribution with enhanced diversity and minimal redundancy, achieving broader coverage of long-tail knowledge across diverse domains. Experimental evaluations at the 3B scale demonstrate that CuraWeb significantly outperforms state-of-the-art baselines, yielding an average performance gain of 1.8% across a wide range of benchmarks, particularly in knowledge-intensive and reasoning tasks.

[AI-185] RE: Training-Free Hallucination Detection for Diffusion Language Models

链接: https://arxiv.org/abs/2607.22661
作者: Pengcheng Weng,Yanyu Qian,Yue Tan,Yixin Liu
类目: Artificial Intelligence (cs.AI)
备注: 25 pages

点击查看摘要

Abstract:Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the hallucination problem. Existing hallucination detection approaches for D-LLMs mainly follow a training-based paradigm, relying on data-driven training to optimize the detector. Such reliance not only limits their generalizability across domains models but also incurs additional training cost and deployment overhead. To address these limitations, we propose TRE, a training-free hallucination detection metric for D-LLMs. TRE is a parameter-free and single-run metric that estimates hallucination risk directly from the entropy signals of a single generation, without requiring any detector training or repeated sampling. TRE extracts entropy signals within the D-LLM decoding process along both the spatial and temporal dimensions. From a token-level spatial perspective, we focus on revealing tokens as the most informative carriers of uncertainty, capturing where uncertainty is actively committed. From a diffusion step-level temporal perspective, we empirically identify the dominance of late-step entropy and hence aggregate these signals with a simple linear weighting scheme to obtain TRE. Extensive experiments on multiple D-LLMs and QA datasets demonstrate that TRE achieves competitive performance, while enjoying strong generalizability, efficiency, and robustness.

[AI-186] StanceBench: A Benchmark for Audio LLM -Based Interpersonal Stance Evaluation from Speech INTERSPEECH2026

链接: https://arxiv.org/abs/2607.22658
作者: Yuzhe Wang(1),Thomas Thebaud(1),Jennifer Hu(2),Jesús Villalba-Lopez(1),Venkatesh Ravichandran(3),Georgi Tinchev(4),Najim Dehak(1),Laureano Moro-Velázquez(1) ((1) Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA, (2) Department of Cognitive Science, Johns Hopkins University, Baltimore, USA, (3) Amazon AGI, USA, (4) Amazon Research, UK)
类目: Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
备注: Accepted to Interspeech 2026

点击查看摘要

Abstract:Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceBench, a benchmark for measuring interpersonal stance in conversational speech and evaluating audio-capable LLMs as automated judges. Using the Seamless Interaction corpus, StanceBench (1) specifies 9 stance dimensions via role-prompt poles, (2) standardizes single-speaker and interaction-based evaluations, and (3) reports LLM-as-a-judge robustness, bias, and stance inference. Across evaluated stances, empathy and politeness are the easiest. Warmth and assertiveness are moderately separable with positivity skew/asymmetry. Honesty is the hardest and shows high prompt order bias, consistent with needing cross-turn evidence. Attentiveness is separable but aligns weakly with humans. Interaction stances are more context-sensitive, with threshold gaps and high variance, especially conflict regulation.

[AI-187] EventOD: Event-Aware OD Flow Generation via LLM -Guided Semantic Modulation

链接: https://arxiv.org/abs/2607.22655
作者: Jie Zhao,Jie Feng,Can Rong,Zhihan Hou,Peng Lu,Yong Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Estimating origin-destination (OD) flows under disruptive events is important for disaster response and urban resilience. Existing deep OD models trained on routine mobility often degrade when extreme events abruptly alter regional functions and population activities, while retraining a new generator for each event is impractical under limited event-time supervision. We propose EventOD, an event-adaptive OD generation framework that steers a pretrained OD generator using structured event semantics. EventOD first uses a large language model to infer region-level functional and demographic control vectors from coarse event observations. It then learns two lightweight adaptation modules, AlphaNet and BetaNet, to calibrate the magnitude of these semantic shifts, and further introduces a retrieval-augmented fallback pathway for scenarios with sparse supervision. The resulting event-conditioned features are injected into a pretrained graph diffusion OD model through input-level modulation, enabling event-aware adaptation without updating generator parameters. Experiments on hurricane- and pandemic-induced mobility across U.S. counties show that EventOD consistently improves both reconstruction accuracy and distributional fidelity over strong baselines. Source code is available at this https URL.

[AI-188] MINT-V2X: A Mobility-Integrated Network Trajectory Dataset for Predictive Resource Management

链接: https://arxiv.org/abs/2607.22654
作者: Abdullah Anjum,Abdolazim Rezaei,Mehdi Sookhak
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Vehicle-to-Everything (V2X) communication systems are based on datasets that not only contain vehicle trajectory data but also wireless network parameters with a realistic level of fidelity, enabling the creation of prediction and optimization models. There is a very critical research infrastructure gap today, and publicly available datasets are likely to be limited to one of the two: mobility or network parameters, and rarely provide a single, integrated view that combines both. This paper introduces MINT-V2X, a comprehensive dataset generated by coupling SUMO traffic dynamics with OMNeT++/Simu5G network simulation. The validation framework is composed of 14 standardized tests based on 3GPP Release 14 (C-V2X), ETSI standards and Shannon capacity theory. The resulting dataset contains 9.87 million synchronized data points from 1,386 vehicles from 15 roadside units (RSUs) during 3 hours of urban traffic simulation. We demonstrate strict algorithmic consistency through network metric correlations (CQI-SINR: 0.993; SINR-PDR: 0.946). Finally, we demonstrate the value of the dataset by conducting an RSU load prediction case study, showing that using trajectory data yields better predictive performance than network-history-only baselines. The dataset, experiments, and complete SUMO configuration files are available in the GitHub repository to facilitate reproduction on alternative simulation stacks.

[AI-189] Do Language Models Converge to Themselves? Recursive Self-Refinement as Textual Relaxation

链接: https://arxiv.org/abs/2607.22653
作者: Xuening Wu,Qianya Xu,Yanlan Kang,Zeping Chen,Yubin Liu,Shenqin Yin
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models are increasingly used in recursive refinement workflows, where an initial draft is repeatedly revised by the same model. Despite their growing use, the long-term dynamics of such workflows remain poorly understood. Does repeated refinement continue to improve outputs indefinitely, or does it converge toward a stable textual form? We study recursive self-refinement as a dynamical process in which repeated LLM revision drives text toward a model-preferred soft fixed-point region. Using GPT-5.5, we generate 10-step refinement trajectories for 50 ICML 2025 abstracts under both default-temperature and deterministic decoding, and additionally evaluate 15 ICML 2020 abstracts. We analyze normalized edit distance, exact and approximate fixed points, word-count stability, exponential relaxation, and external LLM-as-a-judge evaluation. Across all settings, refinement trajectories rapidly saturate. Most edits occur within the first few iterations, after which trajectories enter a soft fixed-point region with only minor surface-level changes. Deterministic decoding reaches exact fixed points earlier and exhibits smaller residual fluctuations than default-temperature decoding, while both achieve universal approximate convergence. The average edit magnitude follows a consistent exponential relaxation pattern, suggesting convergence toward a model-preferred textual equilibrium rather than open-ended optimization. External evaluation indicates that converged abstracts improve clarity, conciseness, and scientific style while preserving technical meaning. These findings support a dynamical-systems view of LLM self-refinement and motivate practical stopping criteria based on edit-magnitude saturation. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22653 [cs.AI] (or arXiv:2607.22653v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.22653 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Xuening Wu [view email] [v1] Fri, 26 Jun 2026 13:44:53 UTC (695 KB)

[AI-190] KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering

链接: https://arxiv.org/abs/2607.22652
作者: Yike Wu,Nan Hu,Guilin Qi,Guohui Xiao,Chen Jiang,Xinchun Zou,Yuchen Lu,Songlin Zhai,Yongrui Chen,Yuyang Zhang,Xiaoguang Li,Lifeng Shang,Jiaoyan Chen,Jeff Z. Pan
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA). Existing approaches primarily combine LLMs with KGs through retrieval-augmented generation (RAG)-based, agent-based, and SPARQL-based methods. Although these methods have achieved notable success, they still suffer from several limitations, including structural information loss, unfaithful reasoning, and limited flexibility and generalization. To address these challenges, this paper proposes KG2Code, a novel approach that transforms knowledge graphs into a code-based representation, preserving structural semantics while naturally aligning with the code-aware pretraining of modern LLMs. Based on KG2Code, KG2Code-QA is further introduced as a KGQA framework that formulates KGQA as a code generation task. This formulation enables the generation of verifiable reasoning traces and executable code, thereby substantially mitigating the impact of hallucinations. In addition, an automated pipeline is developed to construct a large-scale, high-quality code corpus for effectively training open-source LLMs on KG2Code-QA. After training, LLMs are able to perform KGQA in zero-shot scenarios. Extensive experiments demonstrate that the proposed approach significantly outperforms existing KG-enhanced LLM methods for KGQA, while exhibiting strong generalization to unseen KGs. The code and data are available at Github.

[AI-191] ARdena: Scenario-driven control of real-time LLM agents

链接: https://arxiv.org/abs/2607.22651
作者: Luka Borozan,Domagoj Matijević
类目: Artificial Intelligence (cs.AI)
备注: 12 pages, 3 figures

点击查看摘要

Abstract:Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time interactive environments remains a significant challenge. Existing approaches often rely on model fine-tuning or alignment procedures that are difficult to adapt to changing interaction requirements. This paper introduces layered scenario-driven LLM control, a framework that enables runtime behavior control through structured prompting. By combining persistent context with scenario-specific constraints, the approach allows agent behavior to be modified during interaction without changing the underlying model. The framework is implemented in ARDena, a real-time multimodal embodied agent that integrates speech interaction, visual perception, tool use, and avatar-based response generation. The proposed approach is evaluated with respect to control effectiveness, response latency, and operational stability. The results demonstrate that scenario definitions alone can produce substantially different interaction behaviors while maintaining stable real-time operation, highlighting the effectiveness of scenario-driven prompting for controlling LLM agents.

[AI-192] PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving

链接: https://arxiv.org/abs/2607.22648
作者: Meghana Maghyastha,Robert Underwood,Randal Burns,Bogdan Nicolae
类目: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注:

点击查看摘要

Abstract:Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form reusable KV cache prefixes, which are the main technique used by state of art approaches to accelerate inferences. This reduces the latency of accessing the KV cache and alleviates load imbalance caused by a disproportionately large number of requests on servers containing popular tensors. Furthermore, thanks to decentralization, PTStore allows the expansion of the size of the KV cache for LLM inference by orders of magnitude. As a result, PTStore can execute inferences on long passage Q\A datasets 5-6 times more efficiently than current baselines, which do not aggregate memory across different nodes and GPUs and therefore require regenerating the KV cache.

[AI-193] Extracting Algorithms in Pre-trained LLM s: A Case on Hidden Markov Models

链接: https://arxiv.org/abs/2607.22646
作者: Yijia Dai,Zhaolin Gao,Yahya Sattar,Jennifer J. Sun,Sarah Dean
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several candidates without consensus, and none has been grounded in the model’s internal activations. We close this gap with a three-stage pipeline. First, we empirically compare LLM behavior against a suite of candidate algorithms and narrow the space to three classes – though no single class explains LLM behavior across all HMM settings and sequence lengths. Second, we derive theoretical connections between the three classes and show how each can be implemented in-context by a Transformer, validating the construction in a small trained Transformer. Third, returning to pre-trained LLMs, we introduce the Principal Activations Probe (PAP), a layer-wise probing and intervention method that isolates algorithmic signals in model activations. PAP reveals low-dimensional linear representations that causally drive model predictions and track empirical ICL performance. PAP further reveals how these representations shift with properties of the underlying HMM regime; distinct computational stages are localized to different layers. Together, our results connect the in-context behavior of pre-trained LLMs to the underlying internal mechanisms and advance our understanding of how LLMs perform ICL on HMMs.

[AI-194] AutoCluster AutoTopicModeling AutoTrendAnalysis: A Complete AutoML Pipeline for Predicting Emerging Trends

链接: https://arxiv.org/abs/2607.22641
作者: Ahmed Abolfadl,Marwa Mahmoud Abla,Mervat Abu-Elkheir,Maggie Mashaly
类目: Digital Libraries (cs.DL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Conference: 2025 International Conference on Computer and Applications (ICCA)

点击查看摘要

Abstract:Predicting emerging trends is vital for businesses, researchers, and policymakers; yet traditional approaches often lack scalability and adaptability. This paper presents a trend prediction framework based on Automated Machine Learning (AutoML), designed to extract insights from textual datasets with temporal attributes. The system ingests subject-specific textual entries accompanied by a date field. The pipeline begins with preprocessing and embedding, followed by AutoClustering, which uses meta-learning to select the optimal clustering algorithm. AutoTopicModeling then applies successive halving to identify the best topic modeling method: Latent Dirichlet Allocation (LDA), Latent Semantic Analysis (LSA), BERTopic, or Non-negative Matrix Factorization (NMF) based on the coherence score for each cluster. For trend forecasting, AutoTrendAnalysis evaluates multiple models: Facebook Prophet, AutoRegressive Integrated Moving Average (ARIMA), Seasonal-Trend decomposition using Loess (STL), and Long Short-Term Memory (LSTM) selecting the most accurate based on Root Mean Square Error (RMSE), either through successive halving or exhaustive comparison. Topics are classified as strong signals, weak signals, or noise based on forecasting outcomes, enabling the identification of emerging trends. By automating clustering, topic modeling, and time series forecasting, this research enhances trend prediction accuracy while reducing manual effort. The proposed system offers a scalable and user-friendly solution suitable for real-time applications and stakeholders with limited machine learning expertise. Experimental results demonstrate that the proposed system’s best trial achieves a final RMSE of 7.099, indicating high predictive accuracy.

[AI-195] AI-Assisted Causal Inference and Mediation Analyses of Environmental and Psychosocial Determinants of Subjective Cognitive Difficulties in the All of Us Research Program

链接: https://arxiv.org/abs/2607.22640
作者: Cong Cao,Shuangge Ma
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 30 pages, 10 figures

点击查看摘要

Abstract:Short-term environmental exposures have been linked to cognitive and behavioral outcomes, although many reported associations may reflect broader geographic and contextual differences. Using longitudinal data from the All of Us Research Program (2018–2024), we linked daily weather and air-pollution exposures to repeated attention-related and subjective cognitive outcomes. Associations were evaluated using pooled, fixed-effects, lagged, and event-study analyses. Additional machine-learning analyses were conducted to explore potential heterogeneity and latent psychosocial structure. Replication analyses were performed using the 2024 Behavioral Risk Factor Surveillance System (BRFSS). Several environmental exposure measures showed small associations with cognitive outcomes in pooled analyses, but most attenuated substantially after accounting for within-location temporal variation. Mediation, sensitivity, and machine-learning analyses yielded similar conclusions. In contrast, mental-health burden, loneliness, and social functioning were consistently associated with subjective cognitive difficulty and exhibited substantially larger effect sizes than environmental exposures. Similar patterns were observed in BRFSS. Exploratory AI-assisted analyses yielded findings broadly consistent with the primary longitudinal analyses. These findings suggest that short-term environmental perturbations may have limited associations with cognitive outcomes after accounting for within-location variation, whereas psychosocial factors appear to be more consistently associated with subjective cognitive burden.

[AI-196] RACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLM s

链接: https://arxiv.org/abs/2607.22639
作者: Sai Shruthi Sistla,Ashutosh Hathidara,Christopher Toukmaji,Mayank Shrivastava,Karthikeyan Asokkumar
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search. Toolsense shows that this regime has two critical drawbacks: it destroys parametric tool knowledge during training, and its beam-search decoding is too slow for real-time deployment. We introduce TRACE (Tool Retrieval via Augmented Chain-of-thought and Enterprise rules), a two-stage curriculum that resolves this dissociation. Stage 1 reuses the multi-format memorization SFT from ToolSense to seed tool knowledge with LoRA. Stage 2 is our core contribution: the model is trained to emit a thinking trace before producing a JSON list of tool tokens, using two data sources – RRB pairs from ToolSense and queries synthesized to target business rules curated by domain experts – both augmented with reasoning traces. This training objective preserves Stage 1 MCQ and QA probing accuracy while enabling single-beam greedy decoding at production latency. Evaluated on a combined enterprise catalog of 8,300+ tools across two enterprise product lines, TRACE training for Stage 2 not only preserves but improves tool understanding: MCQ accuracy gains +3.2 pp and QA probing gains +9 pp over Stage 1. On retrieval, TRACE achieves ~86% recall on Domain A and ~60% on Domain B – compared to embedding baseline performance of ~27% ~52% – both with single-beam greedy decoding, making it directly deployable at production latency.

[AI-197] Fast Cross-Scenario Adaptation of CSI Models via Channel Conditional Parameter Generation

链接: https://arxiv.org/abs/2607.22637
作者: Xudong Zou,Siyu Wu,Zunlei Feng,Jie Song,Yuanyu Wan,Mingli Song,Jiacong Hu
类目: Artificial Intelligence (cs.AI); Information Theory (cs.IT)
备注:

点击查看摘要

Abstract:Deep learning has shown strong potential for massive multiple-input multiple-output (Massive MIMO) physical-layer tasks, including channel state information (CSI) feedback and channel estimation. However, environmental heterogeneity can severely degrade CSI models in unseen scenarios, while conventional adaptation requires target-domain data and substantial computation. This paper proposes Channel Conditional Parameter Generation (CCPG), an end-to-end pipeline for rapid deployment of CSI models in dynamic wireless environments. CCPG identifies scene-sensitive adaptation bottlenecks through component-freezing experiments and generates only lightweight LoRA weights instead of full model parameters. It compresses high-dimensional channel features into compact latent conditions using cascaded SVD and a Perceiver Resampler. An energy-based canonicalization mechanism mitigates permutation and sign ambiguities in LoRA weights, while a diffusion-based generator incorporates structural information and an asymmetric size-aware loss for topology-aware parameter generation. Experiments on DeepMIMO and WAIR-D for CSI feedback and channel estimation show that CCPG adapts to new scenarios in about 3 seconds with a single forward pass, without target-scenario training or fine-tuning, and achieves cross-domain recovery performance comparable to costly online adaptation. These results demonstrate that CCPG enables efficient deployment of CSI models in large-scale dynamic wireless scenarios for intelligent 6G communications.

[AI-198] Answering Path Queries under Linear and Guarded Existential Rules

链接: https://arxiv.org/abs/2607.22636
作者: Jean-François Baget(LIRMM, Inria, University of Montpellier, CNRS, France),Meghyn Bienvenu(Univ. Bordeaux, CNRS, Bordeaux INP, LaBRI, France),Marie-Laure Mugnier(LIRMM, Inria, University of Montpellier, CNRS, France),Michaël Thomazo(Inria, DIENS, ENS, PSL University, CNRS, France)
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Ontology-mediated query answering is concerned with the problem of answering queries over knowledge bases consisting of a database instance and an ontology. While most work in the area focuses on conjunctive queries (CQs), navigational queries have gained increasing attention. In this paper, we investigate the complexity of answering two-way (conjunctive) regular path queries (©RPQs) over knowledge bases whose ontology is given by a set of guarded existential rules. We first consider the subclass of linear existential rules and show that ©RPQ answering is NL-complete in data complexity, which matches the data complexity of answering RPQs over plain graph databases (i.e., without an ontology). In combined complexity, both tasks are ExpTime-complete in the general case, but RPQ and CRPQ answering drop to PTime-complete and PSpace-complete respectively if there is a bound on predicate arity. For guarded rules, we provide a non-trivial reduction to the linear case, which allows us to show that the complexity of ©RPQ answering is the same as for CQs, namely 2ExpTime-complete in combined complexity (ExpTime-complete in the bounded-arity case) and PTime-complete in data complexity.

[AI-199] CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants

链接: https://arxiv.org/abs/2607.22635
作者: Xuzhao Geng,Haozhao Wang,Xuelian Li,Zhenyu Yang,Haonan Lu,Rui Zhang,Ruixuan Li
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations. However, existing studies are primarily designed for single, explicit goal completion, while phone call assistants face a proxy setting that requires coordinating the device owner’s explicit preset goal with the caller’s implicit and dynamic goal. We introduce \textscCallBench, a Chinese benchmark for evaluating dual-goal coordination in phone call assistants. \textscCallBench contains 50,000 complete multi-turn phone call dialogues across six scenarios: takeout, delivery, taxi, work, life, and harassment. It covers regular presets, emergent presets, and no-preset cases, and includes diverse relations between owner-side and caller-side goals, such as alignment, complementarity, irrelevance, and conflict. We further design a preset-aware turn-level evaluation protocol covering semantic understanding, context use, active guidance, response quality, preset compliance, dialogue rhythm, and safety. Experiments on representative dialogue methods show that existing approaches still struggle with this task, highlighting the need for phone call assistants that can make reliable turn-level decisions between two independent goals under proxy constraints.

[AI-200] PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding

链接: https://arxiv.org/abs/2607.22634
作者: Zheng Wang,Zhifan Ye,Qi Cheng,Yonggan Fu,Ziyan Wang,Feng Zhu,Haozhe Zhao,Jan Kautz,Pavlo Molchanov,Humphrey Shi,Minjia Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, generating tokens in parallel. This makes them effective draft models for speculative decoding (SD), producing an entire block of draft tokens in a single forward pass. Yet existing diffusion-based drafting methods rely on linear drafting, even though dLLMs emit multiple candidate tokens across positions, inducing a large combinatorial space of decoding paths. Consequently, they limit acceptance length and decoding efficiency. To exploit this multi-candidate structure, we apply tree-based drafting to diffusion drafters, enabling exploration of diverse candidate paths. However, we find that naive tree drafting is suboptimal: diffusion marginals are prefix-blind, mismatching the prefix-based AR verification and yielding unreliable path ranking. We propose PRESTO, a principled framework that extends tree-based drafting to diffusion drafters while resolving the fundamental mismatch between diffusion draft confidence and prefix-based AR verification through PREfix-aligned Scoring and priority-based Tree search for diffusion speculative decOding. The key principles behind PRESTO are that (1) candidate ranking should align with the prefix-based nature of AR verification, and (2) tree construction should prioritize candidate paths with high verification potential to maximize acceptance length. Extensive experiments show that PRESTO achieves up to an average of 1.5\times end-to-end throughput speedup on the state-of-the-art dedicated diffusion drafter SD and an average of 1.12\times on self-speculative diffusion LLMs across diverse benchmarks.

[AI-201] Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering

链接: https://arxiv.org/abs/2607.22633
作者: Guixin Su,Qiankun Pi,Mayi Xu,Wenli Li,Ming Zhong,Yuanyuan Zhu,Jiawei Jiang,Tieyun Qian
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Table Question Answering (TableQA) aims to reason over tables to answer user queries. Existing research treats all questions uniformly and evaluates solely through overall accuracy, obscuring a critical reality that LLMs excel at simple lookups yet struggle with complex operations like aggregation and arithmetic. To reveal this disparity, we introduce a novel \emphOperation-wise TableQA task with a fine-grained question taxonomy and release two datasets named WikiTQ-ow and TabFact-ow for evaluation. As for modeling bottlenecks, existing methods flatten tables into linearized texts, disrupting inherent structures and inducing the ``lost-in-the-middle’’ issue, which poses a primary barrier to complex cross-row reasoning. Moreover, they typically reason from scratch, neglecting reusable patterns shared across similar operations. To address these limitations, we propose a Skill-augmented Table Graph Reasoning (SkillTGR) framework for self-evolving structured reasoning. Specifically, SkillTGR represents tables as attributed graphs with explicit row-column-cell structures, where LLMs plan and execute dynamic chains to retrieve evidence subgraphs for graph traversal reasoning. Based on this, SkillTGR builds a hierarchical SkillBank to distill reason trajectories into abstract skills under cognitive heuristics, then hybrid retrieves both successful and failed skills for contrastive augmented table graph reasoning, thereby enabling the continual self-evolution. Extensive experiments demonstrate that SkillTGR achieves superior performance with an average of 5.91% overall and 6.03% operation-wise improvement, also reducing 19.76% token consumption and 27.64% inference latency. Our codes and data will be released upon publication.

[AI-202] VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing ICML2026

链接: https://arxiv.org/abs/2607.22632
作者: Yexiang Liu,Wen Zhong,Sijie Zhu,Xin Gu,Fan Chen,Junxian Duan,Jie Cao,Longyin Wen,Zhenfang Chen
类目: Artificial Intelligence (cs.AI)
备注: ICML 2026

点击查看摘要

Abstract:The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective reward models. To address these challenges, we define a comprehensive vlog evaluation framework guided by professional vlog creators and product managers, establishing a taxonomy of six key dimensions, i.e., Creativity, Consistency, Concept Design, Cinematography, Narration, and Pacing. Subsequently, we curate a large-scale dataset of 100k vlog edits and a dedicated benchmark, VRMBench, to evaluate the vlog rewarding capabilities of Multimodal Large Language Models (MLLMs). Finally, we present VlogReward, a robust vlog reward model that can provide both fine-grained multi-dimensional scores and actionable feedback for iterative refinement. Technically, we enhance the Group Relative Policy Optimization (GRPO) framework by introducing an adjustable inter-group comparison reward, which mitigates the “direction blindness” issue of standard GRPO and enables the model to better distinguish varied-quality edits. VlogReward achieves state-of-the-art results that significantly outperform existing MLLMs, including GPT-5 and Gemini-3-Pro. We hope that our study can help vlog creators and foster automated vlog evaluation and refinement systems.

[AI-203] okenMem: Faithful Knowledge Injection for Frozen LLM s

链接: https://arxiv.org/abs/2607.22625
作者: Chengzhang Yu,Chenyang Zheng,Zening Lu,Yingru He,Yutong Huang,Yiming Zhang,Yue Xu,Zhanpeng Jin
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 15 pages, 3figures

点击查看摘要

Abstract:Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: when retrieved information contradicts parametric memory, the shared self-attention pathway produces unpredictable outputs. We present TokenMem, a lightweight memory system that injects knowledge into frozen LLMs through a dedicated cross-attention channel, bypassing competition with parametric memory in the residual stream. TokenMem trains only a thin gating adapter ( \sim 3-7M parameters) via a two-phase curriculum: first learning general knowledge utilization, then strengthening faithful compliance under counterfactual knowledge. In controlled experiments on five models spanning three families (Qwen3-4B/8B/14B, LLaMA-3.1-8B, OLMo-3-7B), TokenMem achieves 69-70% Knowledge Compliance (KC) on counterfactual benchmarks, compared to 20-52% for vanilla RAG, a gap of up to 49 percentage points. Ablation studies show that the two-phase curriculum is critical: removing Phase 2 collapses KC to near-zero. Mechanistic analysis reveals that the gate adapter learns a conflict-aware, layer-specific injection strategy without explicit supervision.

[AI-204] CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process

链接: https://arxiv.org/abs/2607.22624
作者: Minghao Yang,Yanjun Xu
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Recently, there have been several works in the Text-to-SQL domain that utilize Small Language Models (SLMs) for training. These approaches achieve performance close to that of large models in generating SQL, using only the computational power of a single NVIDIA RTX 4090 GPU, while also ensuring data security. Most existing methods filter out redundant tables and columns during Schema Linking to improve Text-to-SQL accuracy. However, they do not consider the precision-recall trade-off when selecting the candidate schema subset. Our research found that both the precision and recall of Schema Linking directly affect the final SQL accuracy. Therefore, we propose a novel framework for efficiently fine-tuning SLMs on Text-to-SQL tasks, CHS-SQL, that not only balances precision and recall but also improves overall performance on Text-to-SQL tasks. Its main innovation lies in the Schema Linking phase, where a heuristic search combined with model internal confidence is employed to achieve an optimal precision-recall trade-off. This elaborated mechanism maximizes the precision of relevant schema candidates for the generated SQL queries while suppressing irrelevant noise. The same strategy is further applied during SQL generation to refine candidate queries while helping the SLM to avoid trapping in a local optimum. Our method achieves state-of-the-art (SOTA) results on Text-to-SQL tasks via SLMs.

[AI-205] Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning

链接: https://arxiv.org/abs/2607.22621
作者: Aamir Hamid,Bharg Barot,Satvik Racharla,Tim Finin,Primal Pappachan,Roberto Yus
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:While large language models (LLMs) enable strong question answering (QA), budgeted deployment is complicated by nondeterminism and heterogeneous resource profiles (cost, latency, and energy). We present OPTI-Q, a database-inspired, cost-based optimizer that implements a plan-before-execute paradigm for multi-LLM orchestration. OPTI-Q models LLM invocations as physical operators in an execution DAG and, for each question, searches for plans that optimize answer quality (QoA) while trading off financial cost, latency, and energy under user-specified resource constraints. Plans can include sequential operators that pass intermediate answers as context and parallel/blend operators that run models concurrently and merge their outputs. To search this space without executing each candidate plan, OPTI-Q uses PERFDB, a statistics catalog populated and refreshed from benchmarks and execution traces, to estimate the QoA and resource costs of both individual operators and composed subplans. Using these estimates, OPTI-Q performs Pareto-frontier search and selects a final plan based on user preferences. On MMLU-Pro and SimpleQA under user-specified budgets, OPTI-Q improves average QoA by ~58% and ~41% over baselines at comparable cost, demonstrating that database-style planning yields better quality-resource trade-offs for multi-LLM QA.

[AI-206] DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training

链接: https://arxiv.org/abs/2607.22614
作者: Hanlin Du,Zhiyuan Yan,Haiquan Chen,Jiarui Fang,Yungang Bao,Sa wang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.

[AI-207] Quotient Tree Arithmetic: Deferred-Division Computation with Bounded Symbolic Depth and Cross-Subtree Cancellation

链接: https://arxiv.org/abs/2607.22612
作者: Gregory Magarshak
类目: ymbolic Computation (cs.SC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Mathematical Software (cs.MS)
备注: 20 pages, 3 tables

点击查看摘要

Abstract:We introduce Quotient Tree Arithmetic (QTA), a computational substrate in which values are represented as deferred quotient pairs (N, D) whose ratio is evaluated lazily at a designated materialization boundary. The framework applies to any domain: IEEE 754 doubles used as exact integer containers give exact rational arithmetic within the 2^53 exactness window; arbitrary IEEE doubles extend coverage to transcendental values including machine learning activations such as exp(x) and sqrt(x). Three structural theorems underpin QTA. (1) Bounded Depth Growth: each arithmetic operation increases tree depth by at most 1, giving O(m) tree size after m operations with no combinatorial explosion. (2) Cross-Subtree Cancellation: subtrees appearing in both numerator and denominator positions cancel via reference identity without arithmetic, including transcendental values computed once and shared. (3) Deferred Stability: a single IEEE division at the materialization boundary introduces at most one-half ULP of rounding error, versus O(m) ULP for eager evaluation. For machine learning training, QTA provides: structural prevention of gradient underflow to zero; O(1)-cost gradient computation via chain-rule tape collapse when intermediate activations are reference-identical; shared-weight batch compression reducing DAG storage from O(BLd) to O(L+Bd) for a batch of B examples through L layers; and tracked factor cancellation replacing O(log n) GCD with O(1) trial division when denominators are known. We propose a vectorized hardware normalization instruction (RatCleanup) for SIMD-parallel rational pair reduction. The algebraic foundation is the localization of a ring at its multiplicative set, connecting QTA to algebraic structure theory while grounding it in hardware-native IEEE arithmetic. Comments: 20 pages, 3 tables Subjects: Symbolic Computation (cs.SC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Mathematical Software (cs.MS) MSC classes: 65G50, 68W30, 65Y05, 68T07 ACMclasses: G.1.0; B.2.4; I.2.6; F.2.1 Cite as: arXiv:2607.22612 [cs.SC] (or arXiv:2607.22612v1 [cs.SC] for this version) https://doi.org/10.48550/arXiv.2607.22612 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-208] Evaluating LLM s as Interpretable Controllers for Dynamical Systems

链接: https://arxiv.org/abs/2607.22609
作者: Aleksander Østensen,Alberto Mino Calero,Anastasios M. Lekkas,Adil Rasheed
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly used for decision-making and reasoning tasks, yet their potential as controllers for physical systems remains largely unexplored. This work investigates whether LLMs can function as interpretable controllers for a dynamic thermal environment, examining their ability to follow setpoints, interpret natural-language commands, reason about actuator effects, and incorporate prior model-based knowledge. Five LLMs of varying scales are evaluated under multiple scenarios, including settings with penalties on heater or fan usage and cases where the models have access to a physics-based prediction tool. The results show that control performance depends on model complexity: while low- and mid-scale models frequently misinterpret actuator dynamics or generate inconsistent reasoning, high-complexity models such as Qwen-3~14B and GPT-4o achieve accurate temperature tracking, stable actuator usage, and coherent explanations aligned with physical principles. Incorporating a physics-based model significantly improves control smoothness and energy efficiency by enabling anticipatory decision-making. A detailed reasoning taxonomy further reveals a clear progression from causal misinterpretation in smaller models to cohesive and temporally aware reasoning in larger ones. The findings demonstrate that LLMs can act as interpretable controllers when sufficiently capable and appropriately grounded in domain knowledge, highlighting promising opportunities for hybrid model-based and language-driven control strategies that can provide plausible explanations.

[AI-209] Group Preference Collapse in Personalized Multimodal Large Language Models

链接: https://arxiv.org/abs/2607.22603
作者: Fan Lyu,Wenqi Zhang,Joost van de Weijer
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-level information and overlook diverse user preferences. We identify group preference collapse, where multi-user personalized MLLMs become insensitive to individual preferences and drift toward dominant population-level choices due to suppressed preference signals and unreliable preference use during generation. We propose PrefMoE, a preference-centric framework that separates stable profile information from preference-related representations. PrefMoE decomposes preferences into shared prototypes and personalized residuals, preserves individualized residuals with imbalance-aware learning, counterfactual pseudo-user augmentation, and residual decorrelation, and routes profile and preference factors through separate LoRA adaptation paths. Experiments across multiple MLLM backbones show that PrefMoE improves preference-sensitive personalization while substantially reducing preference collapse. Project page: this https URL.

[AI-210] DeepLook: Deeper Thinking with Lookahead

链接: https://arxiv.org/abs/2607.22602
作者: Tingxin Yang,Zefeng Wang,Mengyue Wang,Xingcheng Zhou,Yunpu Ma
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reasoning tasks than parameter scaling alone. However, existing approaches remain inefficient in how compute is allocated within a reasoning trace. Motivated by the observation that reasoning failures often exhibit an early onset of uncertainty before a wrong answer become explicit, we introduce DeepLook, a training-free monitor-and-intervene decoding framework that concentrates lookahead compute at uncertainty bottlenecks. DeepLook aggregates token-level confidence into segment-level signals, triggers when confidence drops relative to recent history, and explores candidate continuations with fixed-horizon lookahead. Branches are ranked by Average Lookahead Confidence (ALC), the average segment-level confidence over rollout continuations, then pruned and aggregated through voting. On four competition-style mathematics benchmarks across DeepSeek-R1-8B, Qwen3-32B, GPT-OSS-20B, and GPT-OSS-120B, DeepLook shifts the accuracy–token-cost Pareto frontier: it improves accuracy over DeepConf-low in 11 of 16 settings while reducing dataset-level token generation by 87.3% on average, including gains of +3.1 on AIME25 with Qwen3-32B and +8.8 on BRUMO25 with GPT-OSS-20B. These results show that selective, future-aware intervention yields substantially stronger accuracy–cost trade-offs than uniformly scaling complete reasoning trajectories. Code is available here.

[AI-211] Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

链接: https://arxiv.org/abs/2607.22600
作者: Ridwan Mahbub,Mohammed Saidul Islam,Md Tahmid Rahman Laskar,Mizanur Rahman,Mir Tafseer Nayeem,Enamul Hoque
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappings-can systematically alter interpretation while preserving the underlying data. As Vision-Language Models (VLMs) are increasingly used for chart understanding and analytical reasoning, assessing their robustness to such deceptive visualizations has become critical for trustworthy data analysis. We introduce VisDeception, the first controlled paired benchmark for evaluating the robustness of VLMs to misleading chart designs. The benchmark contains 1,600 paired faithful and misleading charts spanning eight major categories of deceptive visualization tactics, where each misleading chart is paired with a faithful counterpart generated from the same underlying data. To isolate deception-induced reasoning errors from baseline chart-understanding errors, we introduce the Deception Score, a paired evaluation metric that quantifies how misleading visualizations shift model responses away from the faithful interpretation of the data. Across 32,000 responses from 10 state-of-the-art VLMs, we find that even advanced models remain highly vulnerable to deceptive visual manipulations. To improve robustness, we further propose an inference-time multi-agent mitigation framework that grounds reasoning in structured chart metadata extracted from the visualization before answer generation, enabling models to reduce the influence of deceptive visual cues without requiring explicit user instructions. Together, our findings reveal important reliability gaps in current chart-understanding systems and establish benchmark-driven evaluation, deception-aware metrics, and structured reasoning as promising directions for developing more trustworthy VLMs for visual analytics.

[AI-212] Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting

链接: https://arxiv.org/abs/2607.22599
作者: Chen Su,Yuanhe Tian,Yan Song
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Diffusion models have become a widely used framework for probabilistic time series forecasting, modeling the distribution of future values given an observed history. In time series forecasting, however, the future continues the observed history, creating an asymmetry the standard diffusion process leaves unaddressed, with slowly-varying content largely determined by the observed continuity while higher-frequency dynamics carry most of the residual uncertainty. Existing diffusion-based forecasters decouple this asymmetry through an external rule before generation, leaving the corruption trajectory blind to which parts of the target the history can already anchor. We propose DiffDiff, a diffusion framework that embeds this predictability asymmetry into the diffusion trajectory itself, so that a single end-to-end diffusion process becomes aware of which parts of the target the history can already anchor. DiffDiff makes the forward operator step-dependent so that the noisy intermediate state progressively shifts from the target itself toward its second-order differenced structure, while a conditioning pathway supplies the denoiser with both value-domain and differential history information balanced by a stage-adaptive gate at each diffusion step. The terminal distribution approaches a standard Gaussian, preserving compatibility with existing samplers. On seven benchmarks across four prediction horizons, DiffDiff outperforms six diffusion baselines, and our analysis confirms that DiffDiff concentrates the diffusion’s generative effort on the most uncertain components of the target while relieving it from rebuilding the history-anchored content.

[AI-213] HyCE-RAG : Hypergraph Chain-of-Evidence Retrieval-Augmented Generation for Explainable Multi-hop Question Answering

链接: https://arxiv.org/abs/2607.22597
作者: Hong-Yu An,Yun-Jian Zhang,Chen-Wei Liang,Tian-Yi Zhang,Jian Ding,Yi-Lun Wu,Ao-Bo Li,Wei-Cong Su,Saifullah,Mujiangshan Wang
类目: Artificial Intelligence (cs.AI)
备注: 15 pages, 3 figures, 4 tables

点击查看摘要

Abstract:Multi-hop question answering requires systems to retrieve evidence from multiple documents and connect scattered facts into a coherent reasoning process. Standard retrieval-augmented generation (RAG) mainly relies on semantic similarity between a query and text chunks, and therefore often fails to model structural relations among entities, facts, and evidence units. Graph-based RAG improves this by introducing graph-structured knowledge, but pairwise edges are still limited in representing higher-order associations involving multiple entities and contexts. We propose HyCE-RAG, a Hypergraph Chain-of-Evidence Retrieval-Augmented Generation framework for explainable multi-hop question answering. HyCE-RAG organizes entities, relations, and contextual evidence into hyperedges, builds a query-aware evidence hypergraph, and performs confidence propagation over entity–hyperedge incidence structures. It then uses confidence-guided evidence assembly to select, connect, and rank evidence paths before answer generation. The scoring process jointly considers semantic relevance, entity connectivity, evidence coverage, relation reliability, extraction confidence, and propagated confidence. By providing the language model with structured evidence chains rather than flat retrieved passages, HyCE-RAG supports more faithful and interpretable reasoning. Experiments on HotpotQA, 2WikiMultihopQA, MuSiQue, and two GraphRAG-Bench subsets show that HyCE-RAG consistently outperforms standard RAG and graph-based RAG baselines in answer accuracy, context relevance, and faithfulness. These results suggest that hypergraph-based evidence organization is a promising direction for post-retrieval reasoning in complex question answering.

[AI-214] An Agent ic Orchestration of Atomistic Simulations

链接: https://arxiv.org/abs/2607.22596
作者: Rahul Somasundaram,Adela Habib,Khanh Dang,Sachin Shivakumar,Ryley G. Hill,Golo Wimmer,Avanish Mishra,Aleksandra Pachalieva,Arthur Lui,Hari Viswanathan,Michael Grosskopf,Saryu Fensin,Russell Bent,Nathan DeBardeleben,Earl Lawrence
类目: Artificial Intelligence (cs.AI); Materials Science (cond-mat.mtrl-sci)
备注: 19 pages, 9 figures

点击查看摘要

Abstract:Atomistic simulations are central to materials design, but their execution involves complex, multi-step workflows that require significant human expertise. Here, we present an agent-based system embedded within the URSA (Universal Research and Scientific Agent) framework that automates the design, execution, and validation of atomistic simulations, demonstrated using the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) tool. Our system autonomously selects interatomic potentials, constructs and runs simulations, and performs iterative error recovery within a closed-loop workflow. We evaluate the scientific reliability of the agent by benchmarking its outputs against LAVA, a high-throughput toolkit for LAMMPS and the Vienna Ab initio Simulation Package (VASP) calculations. Our framework reduces manual intervention and trial-and-error, thereby improving the rigor, reproducibility, and scalability of atomistic modeling.

[AI-215] xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps

链接: https://arxiv.org/abs/2607.22595
作者: Michael Blum,Mark Silberstein,Yaniv David
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Mechanistic interpretability (MI) has emerged as a powerful approach for analyzing and intervening in inference computations, with a growing number of applications such as jailbreak attempt detection, truthfulness evaluation, and hallucination detection. Unfortunately, MI deployment in production model-serving systems is currently not practical, as most existing MI frameworks introduce prohibitively high runtime overheads. The fundamental problem is that MI functions do not compose cleanly with served models: they fragment deployment, often force draining requests and rebuilding serving state, and conflict with critical performance optimizations such as continuous batching and CUDA-graph execution, essential for production deployments. We present xMIx, a serving-native framework for deploying MI applications in production inference serving environments. xMIx enables attaching MI functions to a predefined set of locations in the model runtime, interposing on activations within the layers and residual streams. xMIx supports conditional invocation of MI functions depending on the outputs in preceding model layers. Multiple MI applications can be deployed in a single model instance. xMIx compiles them all into the serving path but activates them dynamically at runtime only when necessary, with negligible performance cost, and without requiring a separate model instance or alternative execution stack. We integrate xMIx with the vLLM serving system and evaluate it across three major models and seven diverse MI applications. xMIx achieves performance comparable to native vLLM execution, incurring a slowdown of 1.3% mean inter-token latency (ITL), 1.2% for tail P99 ITL, 2.6% for mean time to first token (TTFT), and 1.6% for mean total token throughput (TTT). Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22595 [cs.AI] (or arXiv:2607.22595v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.22595 Focus to learn more arXiv-issued DOI via DataCite

[AI-216] ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation

链接: https://arxiv.org/abs/2607.22588
作者: Samyak Jhaveri,Erel Kaplan,Tom Yotam,Le Chen,Tomer Bitan,Niranjan Hasabnis,Gal Oren
类目: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 57 pages, 22 figures, 11 tables; includes appendices

点击查看摘要

Abstract:Modern compute-intensive software must migrate across a changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers, including CUDA, OpenMP, OpenCL, and OpenMP target offload. Large language models and autonomous coding agents are increasingly proposed for such migration, but the field lacks reliable ways to measure whether they preserve the low-level parallel semantics that make translations behaviorally valid, including thread indexing, synchronization, memory management, host-device coordination, and API-specific execution structure. We present ParBench, a kernel-centric benchmark framework for evaluating LLM-based parallel API translation under executable, reproducible conditions. ParBench fixes the surrounding build, run, and verification infrastructure through declarative benchmark specifications and asks models to translate only the computational kernels. It draws on multiple open-source HPC suites and covers representative cross-API translation directions among CUDA, OpenMP, OpenCL, and OpenMP target offload. To test whether success reflects robust translation rather than surface-form memorization, ParBench includes AST-driven, intended behavior-preserving, baseline-validated source augmentation. Evaluations on state-of-the-art open and proprietary LLMs show persistent barriers to reliable parallel code translation, including direction asymmetry, multi-file coordination, incomplete API adaptation, and uneven robustness to source-level perturbations. Code is available at this https URL. Comments: 57 pages, 22 figures, 11 tables; includes appendices Subjects: Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC) Cite as: arXiv:2607.22588 [cs.AI] (or arXiv:2607.22588v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.22588 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-217] riSP: Tri-Signal Structured Pruning for Large Language Models

链接: https://arxiv.org/abs/2607.22587
作者: Manel Kara laoua,Soumia Bouyahiaoui,Aicha Boutorh
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute cost of their parameters. Structured pruning addresses this by removing entire structures such as attention heads and Multi-Layer Perceptron (MLP) neurons to produce smaller dense models that run efficiently on standard hardware. However, existing methods rely on either gradient-based importance estimation, which is memory-prohibitive, or activation-based statistical proxies, which do not directly measure the effect of removal on the loss. Furthermore, the interaction between the importance criterion and the post-pruning recovery strategy has not been systematically studied. We propose TriSP (Tri-Signal Structured Pruning), an importance metric that combines weight magnitude scaled by activation norm with first-order gradient sensitivity via a geometric mean, producing a channel-level score that captures both structural and loss-sensitivity signals. Combined with adaptive per-layer budget allocation and low-rank adaptation (LoRA) recovery, TriSP achieves the lowest perplexity and highest zero-shot accuracy across all tested configurations, reaching 6.80 WikiText-2 perplexity at 20% pruning on LLaMA-7B. Inference throughput improves by 82% at 50% pruning, while still maintaining competitive performance.

[AI-218] he Scaffold Effect in Coding Agents : Harness Choice as a Hidden Variable in Coding-Agent Evaluation ICML2026

链接: https://arxiv.org/abs/2607.22585
作者: Naman Vats,Oleg Golev
类目: Artificial Intelligence (cs.AI)
备注: 7 pages, 2 figures, 6 tables. Preliminary work; under review at the 5th DL4C Workshop @ ICML 2026

点击查看摘要

Abstract:Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified. Model-to-model comparison is valid when the harness is fixed; when it varies, performance and efficiency conflate model and scaffold effects. We evaluate Qwen 3.6 Plus and MiniMax M2.5 across three open-source harnesses (Goose, OpenCode, OpenHands-SDK) on a stratified 50-task subset of Terminal-Bench Pro. Harness choice induces up to a 40x difference in tokens per solved task, while paired within-model pass-rate differences remain 0-8 percentage points (95% paired-task bootstrap CIs include zero except for the largest gap). Failure fingerprints replicate across models (REASON for Goose, VERIFY/MAX_TURNS for OpenHands-SDK, idle-loop/TIME for OpenCode), indicating harness-level biases that are largely model-independent. For human-centered coding-agent evaluation, model name alone is an incomplete comparison unit: harness-model pairs determine real-world cost, latency, and oversight burden; no-action turns are a per-task wait tax, not just a token tax. We therefore recommend selecting harness-model pairs by pass rate under token/latency budgets, and reporting token usage, latency, and full harness specifications alongside any model comparison. We release anonymized configs, raw trial logs, aggregated snapshots, and analysis scripts.

[AI-219] Multi-Objective Structured Pruning of LLM s for Latency and Model Size Optimization

链接: https://arxiv.org/abs/2607.22583
作者: Muhammad Junaid Ali,Smail Niar,El-Ghazali Talbi
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of strict latency, memory, and energy constraints. Their large parameter counts and computational demands hinder efficient execution on resource-constrained platforms. Although model pruning has emerged as a viable solution for reducing scale while preserving performance, jointly optimizing layers, attention heads, and Multi-Layer Perceptron (MLP) dimensions remains highly complex. Exhaustively exploring this combined design space is computationally expensive and often leads to local optima or unstable configurations. To address these limitations, we propose a hardware-aware, multi-objective structured pruning framework. The proposed two-stage method explicitly targets latency and model size for efficient deployment on edge devices. In the coarse-grained stage, multi-objective depth pruning removes entire attention and MLP blocks to reduce computational load and memory usage. In the subsequent fine-grained stage, Parallel Bayesian Optimization (PBO) searches for the optimal layer-wise pruning ratios for pruning under latency constraints, while importance-based strategies rank the specific components to be pruned within each layer’s allocated budget. Experimental results show that our approach reduces model complexity with minimal impact on commonsense reasoning tasks and zero-shot performance. Our method achieves a favorable trade-off among accuracy, latency, and model size, making it suitable for edge deployment. Across multiple LLMs at 37.5% and 50% pruning ratios, the proposed approach achieves better performance on commonsense reasoning tasks than existing methods while significantly reducing inference cost.

[AI-220] HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization ICML2026

链接: https://arxiv.org/abs/2607.22578
作者: Size Li,Zhiqing Tang,Hongrui Liang,Jianxiong Guo,Jiong Lou,Tian Wang,Weijia Jia
类目: Artificial Intelligence (cs.AI)
备注: to be published in ICML 2026

点击查看摘要

Abstract:The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. However, existing solutions typically prioritize intra-workflow optimization, largely neglecting the significant potential for inter-workflow optimization. In this paper, we propose HeraSys, an LLM serving system designed to optimize the end-to-end performance of concurrent workflows. Through fine-grained orchestration, HeraSys eliminates cross-workflow computational redundancy via structural node merging and reuse. Furthermore, HeraSys introduces a load-aware joint scheduling policy that dynamically manages execution order by evaluating both inter- and intra-query priorities. By integrating a resource skewing mechanism with adaptive batching and pipeline decomposition, HeraSys effectively mitigates tail latency while maintaining low average latency, thereby substantially improving system throughput. Extensive experiments demonstrate that HeraSys reduces P99 latency by up to 2.17 \times and increases serving throughput by up to 1.85 \times under strict latency guarantees.

[AI-221] cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLM s

链接: https://arxiv.org/abs/2607.22577
作者: Xin Yang,Yemin Wang,Mingda Liu,Letian Li,Shuaishuai Cao,Zhengxiao He,Ryan Dong
类目: Artificial Intelligence (cs.AI)
备注: 15 pages, 3 figures

点击查看摘要

Abstract:Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck as models approach trillion-parameter regimes. We aim to scale capacity through MoE-style mixture throughout the LLM pipeline rather than only the FFN. Prior pipeline-level approaches include ParaScale, which introduces virtual tokens and parallel streams but incurs substantial overhead and suffers from homogenized routing and gradient collapse, and AltUp, which uses an auxiliary prediction branch but offers limited adaptivity and slow convergence. We establish that MoE-style mixture layers can be reformulated as variable-kernel dynamic convolutions, where each expert corresponds to a 1\times1 convolutional kernel and routing implements input-conditioned kernel aggregation. Building on this equivalence, we introduce cMoLLM: a convolutionally gated mixture-of-LLMs that routes over end-to-end streams through fully differentiable dynamic convolution. In GPT-2-style models trained on FineWeb, cMoLLM improves language modeling perplexity and downstream GLUE and SQuAD accuracy under matched compute, with better stream utilization, more stable optimization, and favorable scaling compared to ParaScale- and AltUp-style baselines.

[AI-222] mporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models

链接: https://arxiv.org/abs/2607.22575
作者: Mathis Pink,Vy Ai Vo,Qinyuan Wu,Jianing Mu,Javier Turek,Uri Hasson,Kenneth A. Norman,Sebastian Michelmann,Alexander Huth,Mariya Toneva
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Human episodic memory supports the retrieval of experiences that unfold over extended timescales, yet the computational mechanisms underlying this ability remain debated due to the limited mechanistic accessibility in long-term memory experiments in humans. Long-context LLMs may offer promising ways to reveal plausible computational mechanisms that drive this type of retrieval. Here, we investigate whether and how LLMs capture the core behavioral signatures of episodic memory via a temporal order memory task. Using a new dataset of human behavior based on memory of a full-length novel, we show that models exhibit the same characteristic distance effect observed in humans on this task. We next apply long-context mechanistic interpretability analyses to uncover how models solve this task, and find that model performance relies on a one-dimensional temporal code that is reinstated during retrieval by a single time-reinstatement attention head. These findings support temporal context reinstatement as an important mechanism for episodic-like temporal-order memory in LLMs, offering new insights into how temporal aspects of long-term episodic memory may be instantiated in both artificial and biological systems.

[AI-223] PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability

链接: https://arxiv.org/abs/2607.22573
作者: Wen-Kao Li,Ze-Feng Gao,Zhong-Yi Lu
类目: Artificial Intelligence (cs.AI); Materials Science (cond-mat.mtrl-sci); Computational Physics (physics.comp-ph)
备注: 18 pages, 3 figures, includes Supplementary Information

点击查看摘要

Abstract:Imaginary phonon modes remain a practical bottleneck in computational materials screening because otherwise plausible structures can be locally dynamically unstable under a chosen workflow. Here we present PhononBench-MP40, a spectrum-resolved benchmark dataset of Materials Project-derived crystals for workflow-defined phonon stability. The dataset starts from 47,969 MP40 workflow tasks and provides 46,899 completed records with paired stability labels and local phonopy YAML spectra, including 16,683 Stable records and 30,216 completed-phonon unstable records. A further 1,067 relaxation failures are reported separately rather than merged into the completed phonon denominator. The release centers on the local YAML spectrum: the stability label, the lowest sampled frequency and any threshold-dependent relabeling are derived from that spectrum. The dataset is openly available through Science Data Bank at this https URL. A companion GitHub repository provides the calculation code and lightweight access utilities. PhononBench-MP40 provides an auditable reference for workflow-defined stability classification, minimum-frequency analysis, threshold studies and failure-aware triage, while keeping the reference workflow, data schema and interpretation boundaries explicit.

[AI-224] Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL

链接: https://arxiv.org/abs/2607.22572
作者: Sanjay Mishra,Divya Chukkapalli,Ganesh R. Naik
类目: Artificial Intelligence (cs.AI)
备注: 18 pages , 3 figures , 14 tables ,1 algorithm

点击查看摘要

Abstract:Large language models can generate fluent SQL from natural language, but on real enterprise Oracle databases they frequently fail at execution time: columns and aliases are hallucinated and dialect-specific syntax is missed, leading to ORA-00904 invalid-identifier errors. In this setting, failures are primarily due to missing schema grounding: the model cannot know which tables and columns actually exist. This paper introduces Schema-Aware Localisation (SAL), a lightweight middleware layer for Oracle NL2SQL that requires no model retraining. SAL queries Oracle’s USER_TAB_COLUMNS catalog to build a live schema map, selects a relevant table subset for each question (falling back to the full schema for multi-table queries), and injects this ground-truth context into the LLM prompt. Generated SQL is then checked by the Hallucination Index (Hidx), which validates every this http URL reference against the live catalog, automatically rewrites predictable prefix errors, and otherwise triggers a structured retry with itemised corrections. We evaluate SAL on 500 TPC-H natural language questions executed against a live Oracle Autonomous Database 23c instance using GPT-4o-mini. Without any schema grounding, execution-grounded truth (EGT; executes and matches the reference result set) is 2.2% (12/500). A hand-written static schema hint brings EGT to 62.0%. SAL, with no manual schema curation, achieves 62.6% EGT (96% simple, 95% medium, 40.7% complex) while reducing execution failures from 97.6% to 2.6%.

[AI-225] SCAIR: Schema-Conditioned Agent ic Iterative Reasoning for Enterprise Knowledge Graphs ACL2026

链接: https://arxiv.org/abs/2607.22571
作者: Prateek Chaturvedi,Yuqicheng Zhu,Hongkuan Zhou,Dongzhuoran Zhou,Yunjie He,Steffen Staab,Fei Du,Jie Tang,Evgeny Kharlamov
类目: Artificial Intelligence (cs.AI)
备注: Accepted at ACL 2026 Industry Track

点击查看摘要

Abstract:Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, yet existing agentic approaches that perform well on public benchmarks often fail to generalize to real-world enterprise Knowledge Graphs (KGs), which are dense, schema-driven, and operationally constrained. To address these limitations, we propose SCAIR (Schema-Conditioned Agentic Iterative Reasoning), a training-free framework that integrates structured planning with controlled iterative reasoning by injecting schema-conditioned structural priors and enforcing schema-aware traversal during multi-hop reasoning. Experiments on an enterprise-oriented benchmark constructed from a real-world Configuration Management DataBase (CMDB) demonstrate that SCAIR substantially improves performance over existing KG-RAG methods. Crucially, our study highlights that reliable enterprise graph reasoning cannot rely on generic agentic designs; instead, it must explicitly incorporate the target domain’s structural and operational constraints into the reasoning process. We demonstrate that by aligning agent design with business logic, substantial performance gains can be achieved without the need for costly model retraining.

[AI-226] Reference Feature Atlases for Mechanistic Auditing of Language Models

链接: https://arxiv.org/abs/2607.22570
作者: Rui Wu,Tong Che
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained once on a reference panel and reused for new targets, which attach by fitting only a linear decoder. This yields two complementary views. The atlas channel reads the target on already interpreted panel features, providing a stable coordinate system across models. The residual channel learns features only from what the atlas fails to reconstruct, making “outside the reference panel” an explicit audit signal. We train leave-one-out atlases over five 7-9B instruction-tuned models and audit held-out Mistral and Qwen targets. On three controlled LoRA hidden objectives injected into both targets, the residual channel makes the planted mechanism perfectly controllable at runtime while matched controls stay unaffected and recovers the planted objective as the top-ranked latent across both targets; on Mistral, where the per-target SAE and pairwise crosscoder baselines are retrained for a head-to-head benchmark, both baselines fail to do so. On Qwen-2.5, the same channel additionally reveals a panel-relative political-framing cluster; steering it shifts the audited framing metrics while out-of-domain controls remain unchanged. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22570 [cs.AI] (or arXiv:2607.22570v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.22570 Focus to learn more arXiv-issued DOI via DataCite

[AI-227] Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines

链接: https://arxiv.org/abs/2607.22569
作者: Yifei Ge,Weisong Sun,Jinkun Xiao,Yuchen Chen,Yebo Feng,Peizhuo Lv,Xia Feng,Chunrong Fang,Zhihong Zhao,Zhenyu Chen,Yang Liu
类目: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注: Preprint. 12 pages, 6 figures

点击查看摘要

Abstract:Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution environments, and the underlying system. For example, if a coding agent inserts a hook into a system startup or configuration script, that change can persist after the interaction, be triggered later, and abuse delegated user or system privileges to modify the system. This makes security testing a system problem: the key question is not only what the agent says, but what it actually does to the surrounding environment. We present an execution-grounded red-team testing framework for probing this execution-layer security boundary using observable sandbox evidence, including tool invocations, runtime traces, and file-system diffs. Our framework embeds target unsafe operations into routine software engineering workloads, including unit testing, regression testing, crash reproduction, and validation, and uses an execution oracle to guide refinement when an initial probe is rejected or fails. Across multiple agent frameworks and model backbones, our red-team workload reformulation substantially increases verified unsafe execution, reaching 73.61% on code carriers and 53.93% on text carriers. These results show that coding agents in system operations remain insecure under task disguise: once risky intent is hidden inside plausible engineering tasks, the agent can be induced to carry out unsafe actions on the surrounding system. More broadly, coding agents in system operations still demand stronger security testing and safeguards.

[AI-228] Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting

链接: https://arxiv.org/abs/2607.22568
作者: Ruiyi Tao,Xiaolong Tu,Haoxin Wang
类目: Artificial Intelligence (cs.AI); Networking and Internet Architecture (cs.NI)
备注:

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly deployed on mobile and embedded devices to improve privacy and reduce network latency. Yet on-device inference faces a fundamental constraint: high energy consumption on battery-powered, resource-limited hardware. While model compression and runtime acceleration have been widely studied, the effect of \emphprompt design on energy efficiency remains underexplored. This paper presents an empirical study of the relationship between prompt wording and energy consumption for on-device LLMs. Using real power measurements collected on a smartphone, we quantify how linguistic features, particularly imperative keywords and instruction structure, affect decoding length and total energy. Our results show consistent energy differences across verbs and tasks, indicating that prompt engineering is a lightweight lever for improving energy efficiency.

[AI-229] MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models

链接: https://arxiv.org/abs/2607.22566
作者: Zeyu Zhang,Ziqing Wang,Kaize Ding
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue. Existing medical QA benchmarks largely test short context knowledge or single document grounding, leaving open whether LLMs can use, connect, and abstain over longitudinal patient histories. We build MedLoCoMo from deidentified MIMIC-IV and MIMIC-IV-Note records by constructing admission-level clinical packets, synthesizing grounded doctor-patient conversations, and generating evidence linked QA items over single-admission, cross-admission, and adversarial unanswerable settings. The benchmark contains 100 patient timelines averaging 1,669.8 turns, 29.7 sessions, and 74,512.2 tokens per conversation. Across the evaluated baselines, cross-admission reasoning is consistently harder than localized evidence use, even when models have long context windows or use external memory or retrieval methods. The code and MedLoCoMo benchmark release is available at this https URL for use and reproducibility.

[AI-230] DSTFView: Multi-View Cloud-Edge Workload Forecasting with Dual-Input Spatio-Temporal-Frequency Modeling

链接: https://arxiv.org/abs/2607.22565
作者: Qingzhong Li,Hui Ma,Yajun Zhang,Qingchang Ma,Zhou Long
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted in WASA 2026

点击查看摘要

Abstract:With the widespread deployment of edge-side AI inference, edge platforms are increasingly required to support latency-sensitive, highly concurrent, and reliability-critical applications. However, existing methods often struggle to balance multidimensional feature modeling and forecasting efficiency in collaborative cloud-edge environments. To address this issue, we propose DSTFView, a dual-input spatio-temporal-frequency multi-view workload forecasting framework for collaborative cloud-edge environments. It jointly models closeness and period dependencies and extracts spatial, temporal, and frequency-domain dependencies. Besides, it designs an adaptive fusion mechanism and adjusts the contribution of each view to capture abrupt changes. Experimental results on the CPU and TP datasets demonstrate that DSTFView consistently outperforms representative baselines across multiple forecasting horizons and evaluation metrics.

[AI-231] Loss-Aware Feature-Map Pruning in Convolutional Neural Networks Using Multi-Armed Bandits

链接: https://arxiv.org/abs/2607.22564
作者: Salem Ameen,Sunil Vadera
类目: Artificial Intelligence (cs.AI)
备注: 18 pages, 4 figures

点击查看摘要

Abstract:Convolutional neural networks often contain redundant feature maps that increase storage and inference cost. This paper presents a loss-aware feature-map pruning framework using multi-armed bandits. Feature-map pruning is structured because it removes complete convolutional output channels and their producing filters rather than isolated scalar weights. Each candidate feature map is treated as an arm. At each play time, one map is temporarily masked and evaluated on a sampled mini-batch; the map is then restored and the observed loss change is converted into a safe-removal reward. After a fixed play budget, candidate maps are ranked by learned scores and the top-k maps are permanently removed with their filters, biases and corresponding next-layer input-channel kernels. The study evaluates UCB1 and Thompson Sampling, compares them with direct/oracle-style evaluation on LeNet/MNIST, and extends the evaluation to MNIST, CIFAR-10, CIFAR-100, SVHN, CUB-200-2011 and Oxford Flowers 102. Results show that UCB1 and Thompson Sampling preserve accuracy close to unpruned models while removing feature maps and reducing convolutional computation. Friedman and Nemenyi tests show that UCB1 obtains the highest mean rank, followed by Thompson Sampling; both significantly outperform greedy and magnitude-based pruning while remaining statistically comparable to the original unpruned model.

[AI-232] Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

链接: https://arxiv.org/abs/2607.22563
作者: Sagar Chethan Kumar,Rohith Kanathur,Dhaval Patel,Kaoutar El Maghraoui
类目: Artificial Intelligence (cs.AI)
备注: 19 pages, 3 appendices

点击查看摘要

Abstract:Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks such as AssetOpsBench rely on manually authored scenarios and cover a limited set of asset classes. We extend AssetOpsBench with a Smart Grid Transformer asset class and four IEC-grounded diagnostic tools for health-index prediction, dissolved-gas analysis, winding-temperature assessment, and load-profile assessment. We further introduce ScenarioGeneratorAgent, a pipeline for synthetic industrial-agent scenario generation. The pipeline constructs evidence-grounded asset profiles, allocates coverage-aware scenario budgets across operational domains, and generates candidates through a hybrid validation-and-repair loop that enforces schema validity, tool reachability, physical plausibility, standards alignment, and deduplication. To improve scalability, we apply two-level caching, parallel focus-group generation, thread-pool offloading, batched LLM calls, and early rejection filtering. On Smart Grid Transformer scenario generation, these optimizations reduce end-to-end runtime by 8\times for 50 scenarios while preserving quality, achieving a composite quality score of 74.2 \pm 1.9 compared with 73.8 \pm 3.0 for the unoptimized baseline. These results show that standards-grounded synthetic scenario generation can efficiently expand industrial-agent benchmarks without sacrificing scenario quality.

[AI-233] SF-AMS: Strategic Forgetting for Structured Memory in LLM Agent

链接: https://arxiv.org/abs/2607.22562
作者: Ning Yang,Siqi Li,Miaoxin Shen,Yuan Zhou,Meng Zhang,Tong Li,Haijun Zhang
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Managing long-context dependencies remains a primary bottleneck in LLM agents, as redundant and irrelevant information can degrade multi-step reasoning. Strategic Forgetting for Agent Memory Systems (SF-AMS) is proposed as a framework for maintaining compact high-utility memory by modeling the long-term importance of memory units. SF-AMS replaces static retrieval and heuristic decay with a utility-driven survival mechanism that updates memory importance from usage redundancy and temporal signals, inducing a hierarchical memory structure that prioritizes stable entity-consistent information while filtering noise. On top of this, Composite Importance Scoring integrates semantic and entity level signals to improve retrieval robustness. Experiments on LoCoMo and LongMemEval-s show consistent gains over strong state of the art baselines including LightMem MemO and A-Mem. The largest improvement appears in multi-hop reasoning under Qwen2.5-7B where SF-AMS achieves plus 9.65 F1 over the strongest baseline followed by temporal reasoning under GPT-4o-mini plus 6.91 F1 and open-domain tasks plus 6.53 F1 demonstrating strong cross backbone generalization. These results show that modeling memory importance as a dynamic utility signal is critical for reliable long-context reasoning.

[AI-234] Codifying the Judge: Scalable Evaluation via Program Distillation

链接: https://arxiv.org/abs/2607.22561
作者: Tzu-Heng Huang,Shengqi Qiu,Frederic Sala
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions – limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs’ verdicts outperforms one trained on a proprietary LLM’s labels at two orders of magnitude lower API cost.

[AI-235] MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models

链接: https://arxiv.org/abs/2607.22556
作者: Dong Li,Yanchi Liu,Xujiang Zhao,Wei Cheng,Zhengzhang Chen,Xintao Wu,Zhong Chen,Chen Zhao,Haifeng Chen
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Continual learning (CL) is essential for small language models (SLMs) to adapt to evolving real-world needs in resource-constrained deployments. However, directly updating their limited parameter space causes catastrophic forgetting. While memory-based methods naturally address this by decoupling knowledge retention from parameters, existing approaches designed for large language models (LLMs) rely on abundant storage and strong in-context reasoning that SLMs lack. To address these challenges, we propose MIITA, a Memory-Induced Inference-Time Adaptation framework for supervised CL under constrained storage. MIITA stores supervised experiences as compact correction-direction prototypes with semantic anchors, and retrieves them at inference time using semantic and uncertainty-based cues. The retrieved directions are applied through gated temporary hidden-state adaptation, enabling non-destructive reuse of past supervision without backbone updates, prompt extensions, or test-time backpropagation. A local theoretical analysis links this design to first-order loss reduction, uncertainty-guided retrieval, and directional coverage for retaining old-stage knowledge. Extensive experiments across diverse supervised CL settings show that MIITA consistently improves final performance and mitigates forgetting under fixed memory budgets.

[AI-236] DeepLens Diagnosis Agent : Agent ic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLM s

链接: https://arxiv.org/abs/2607.22555
作者: Mahmood Bayeshi,Veysel Kocaman,Muhammed Ali Naqvi,Yigit Gul,David Talby
类目: Artificial Intelligence (cs.AI)
备注: 20 pages, 5 figures, 5 tables. Technical report, John Snow Labs

点击查看摘要

Abstract:Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagnostic reasoning. We present the DeepLens Diagnosis Agent, a five-stage harnessing pipeline (combining model capabilities with disciplined process constraints) centered on a small medical reasoning model (JSL Medical Small 7B v2) and retrieval-augmented generation (RAG). The pipeline enforces structured clinical extraction, disciplined retrieval, constrained candidate generation, explicit evidence triangulation, and an auditable final decision. On the 915-case DiagnosisArena benchmark, the agent achieved 60.14% top-1 diagnostic accuracy, the highest among small and medium-sized models. The same model without the agent workflow achieved 23.99%, a +36-point gain from workflow design alone, despite 88.2% on standard medical benchmarks, showing that diagnostic reasoning under uncertainty requires more than knowledge recall. The agent costs USD 0.0072 per case (24K tokens on A100) with 24-second latency, 35-45% cheaper than Claude Sonnet 4.5 (USD 0.0110) and Gemini 3.1 Pro (USD 0.0128) while outperforming them by +9.70pp and +9.17pp. Harnessing can also correct frontier model failures; workflow constraints can outweigh parameter count or API cost. Beyond aggregate accuracy, the pipeline produces structured intermediate artifacts that make each stage inspectable and support error localization. These properties support high-stakes settings where traceability, reproducibility, and auditable evidence matter alongside benchmark performance. Comments: 20 pages, 5 figures, 5 tables. Technical report, John Snow Labs Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22555 [cs.AI] (or arXiv:2607.22555v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.22555 Focus to learn more arXiv-issued DOI via DataCite

[AI-237] QFoldAgent : An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction

链接: https://arxiv.org/abs/2607.22549
作者: Winson Chen,Yuqi Zhang,Sixu Chen,Nuo Xu,Qiang Guan,Caiwen Ding
类目: Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Hybrid quantum-classical protein structure prediction depends strongly on Hamiltonian penalty weights, yet existing lattice-based workflows typically fix these coefficients by hand and evaluate only very short fragments in simulation. We present QFoldAgent, a closed-loop multi-agent framework for 5-residue tetrahedral-lattice folding in which a design agent proposes sequence-conditioned penalties, a VQE-based quantum-classical pipeline optimizes the resulting Hamiltonian under Qiskit Aer noise, and a feedback agent uses energy-landscape diagnostics and MolProbity validation signals to refine penalties across cycles. Ground-truth metrics such as RMSD are never exposed to the agents and are used only for evaluation. We study the framework on two complementary datasets: 55 QDockBank-derived fragments with known structures and 100 coverage-optimized unseen sequences. On the QDockBank benchmark, QFoldAgent reduces median RMSD from 3.64 Å to 3.20 Å, with the largest gains on the hardest targets. On unseen sequences, the closed loop raises structural validity from 87.5% to 98.7%, recovers 87% of initially invalid cases, and the strongest controller improves cycle-3 energy on 87% of sequences while maintaining 96% Ramachandran-favored geometry. These results show that iterative agent control can systematically improve optimization behavior and reduce failure cases in a 5-residue quantum setting.

[AI-238] SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series

链接: https://arxiv.org/abs/2607.22548
作者: Giovanni B. Esposito,Francesco Antici,Daniele Cesarini,Andrea Bartolini
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Performance (cs.PF)
备注:

点击查看摘要

Abstract:Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, environmental parameters, and physical metrics. Current machine learning approaches for HPC and its telemetry typically rely on a static subset of anonymous, fixed-position sensor variables tailored to single tasks. Consequently, these models become obsolete when target tasks change or sensor metrics vary. We propose SeT-Diff, the first foundational model for compute node telemetry and time-series. Unlike rigid architectures, our diffusion-based approach conditions the generative process on each sensor’s semantic description, decoupling the system dynamics from the structure of the dataset. Experiments on a real-world supercomputer dataset demonstrate a Mean Absolute Error (MAE) of 0.0470 on reconstruction tasks. SeT-Diff exhibits zero-shot permutation stability, maintaining accuracy with negligible degradation even when sensors are shuffled. A single pre-trained model effectively performs data imputation, forecasting, and virtual sensing - achieving a 0.033 MAE in thermal inference - making SeT-Diff an effective data-driven digital twin for HPC systems.

[AI-239] Evaluating Large Language Models for Symbolic Security Protocol Analysis

链接: https://arxiv.org/abs/2607.20712
作者: Paolo Modesti,Syed Ahmed,Ioannis Sfyrakis,Derek Enodolomwanyi
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 36 pages, 5 figures

点击查看摘要

Abstract:Security protocol verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether Large Language Models (LLMs) can perform comparable analysis. We test GPT and DeepSeek in chat and reasoning modes over three runs on 130 obfuscated AnB/AnBx protocols covering 388 security goals, scored against ProVerif and OFMC. Chat models reach 69 to 81% recall at precision below 31%. Reasoning models reverse this trade-off, reaching 66.5% precision for GPT and 45.4% for DeepSeek, but detect just over half the attacks. DeepSeek’s two modes share one underlying model, so the comparison isolates reasoning itself, which raises precision from 27.2% to 45.4%. The GPT contrast spans a model-version change and is only suggestive. All models perform worst on authentication goals: reasoning models detect well under half of injective and non-injective agreement attacks, whereas chat models over-flag them at low precision. Confidentiality is the exception, with F1 up to 95.7% in reasoning mode. Verdicts are unstable across runs, identical on 89.7% of goals for GPT but 74.0% for DeepSeek. Self-reported confidence is uniformly high yet shows no meaningful correlation with correctness. On this benchmark LLMs do not match formal verification, but may serve, at best, as pre-screening filters.

[AI-240] Efficient LLM -Generated Shuttling Compilers for Complex Trapped-Ion Architectures

链接: https://arxiv.org/abs/2607.24714
作者: Fabian Kreppel,Reza Salkhordeh,Ferdinand Schmidt-Kaler,André Brinkmann
类目: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)
备注: 56 pages, 6 figures, 7 tables

点击查看摘要

Abstract:Trapped-ion quantum computers rely on shuttling compilers, which cast an input algorithm into a sequence of ion-qubit movements within a given architecture. We present the first study in which a single frontier large language model (LLM), Claude Opus 4.7, generates and iteratively refines the full Python code of shuttling compilers from written specifications. We start with a compiler for (i) a linear segmented trap, extend it to (ii) a trap with junctions, and finally achieve efficient compilation for (iii) a broad class of connected trap graphs. The compilers for the more general cases are seeded with code from the previous ones. We benchmark the LLM-generated compilers against state-of-the-art hand-crafted ones using a common suite of quantum circuits. The number of shuttling timesteps is reduced by up to 76% for (i) and up to 39% for (ii). For the broad case (iii) of freely connected architectures, we find large variations in the required number of shuttling timesteps, depending on the connectivity. A densely connected, junction-rich architecture yields an order-of-magnitude reduction in shuttling timesteps compared to a corridor-like one. Repeating the complete generation and evaluation with a second frontier LLM, Claude Fable 5, reproduces these findings, with the Fable 5 compilers surpassing the hand-crafted ones more often on the largest circuits. Our results show that an unmodified frontier LLM can produce working, correct, and competitive shuttling compilers without additional manual algorithmic engineering, thus reducing the development time for new architectures from several months to a few days.

[AI-241] Multivariate Time Series Forecasting with Adaptive Non-Local Observables

链接: https://arxiv.org/abs/2607.24399
作者: Yu-Ting Lee,Huan-Hsin Tseng,Samuel Yen-Chi Chen
类目: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Multivariate time series forecasting (MTSF) predicts future values of multiple variables from historical data. While quantum neural networks have been increasingly applied to this task, they typically rely on fixed local measurements, which restrict their expressivity. We propose MTSF-ANO, a simple hybrid model for MTSF that integrates variational quantum circuits with adaptive non-local observables (ANO). On the four ETT datasets, MTSF-ANO ranks first or second in MSE in 17 of 20 settings, improving over the strongest baseline by up to 20% on ETTh1, and outperforms or matches its fixed local observable counterpart across all settings. Our ablations show how the quantum circuit design and ANO non-locality affect performance. These results suggest that ANO is a promising direction for quantum time series forecasting.

[AI-242] acher Knows It Best: Spontaneous Symmetry Breaking and Tipping Points in Networked Langevin Dynamics AI Sycophancy

链接: https://arxiv.org/abs/2607.24304
作者: Sayantari Ghosh,Saumik Bhattacharya,Partha Pratim Chakrabarti
类目: Physics and Society (physics.soc-ph); Artificial Intelligence (cs.AI); Dynamical Systems (math.DS)
备注:

点击查看摘要

Abstract:We formulate a statistical physics framework to model a networked stochastic dynamical system exhibiting bistability, driven by additive noise and social conformity. We apply this model to understand and mitigate AI-induced delusional spiraling-a phenomenon where algorithmic sycophancy from Large Language Models continuously reinforces inaccurate beliefs within a socially interacting society. By partitioning the network into a majority of regular agents and a minority of “aware” nodes (Teachers) placed at topological hubs, we use a degree-weighted mean-field approximation to reduce high-dimensional coupled Langevin equations into a single macroscopic drift equation. We provide a closed-form analytical derivation for the deterministic critical tipping time through a saddle-node bifurcation. We validate this analytical boundary using finite-size scaling and demonstrate a universal data collapse across diverse network topologies. Finally, we optimize an intervention strategy under a strict budget constraint that balances the topological footprint against driving velocity. We prove mathematically that under certain conditions, a highly concentrated, rapid intervention targeting massive hubs strictly outperforms a distributed, slow approach to rescue the network.

[AI-243] HELIOS: An LLM -Driven Autonomous Indirect Trajectory Optimization Agent

链接: https://arxiv.org/abs/2607.24051
作者: An-yi Huang
类目: Instrumentation and Methods for Astrophysics (astro-ph.IM); Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注:

点击查看摘要

Abstract:Low-thrust trajectory optimization is a core technology in deep-space mission design. Indirect methods based on Pontryagin’s Minimum Principle (PMP) offer rigorous optimality guarantees, yet their practical application faces three bottlenecks: (1) transversality conditions must be derived case by case for each constraint type; (2) different dynamics models require repeated code rewrites; and (3) shooting equations are highly sensitive to initial guesses. This paper presents HELIOS (Heuristic Engine for Low-thrust Interplanetary Optimization System), a trajectory optimization agent built around a large language model (LLM). Given a physical problem described in natural language, the system autonomously performs PMP symbolic derivation, SymPy verification, C++ shooting-code generation, and numerical solution without human intervention. Key innovations include: (1) a constraint-adaptive derivation framework that unifies arbitrary constraints into psi(x,p)=0 form and automatically generates stationarity conditions for free parameters (e.g., gravity-assist turning angle); (2) dynamics-adaptive four-module code generation supporting non-standard dynamics (solar sail, J2 perturbation) without modifying the underlying template; and (3) a general derivation rule set covering critical error-prone points in PMP derivation. Experiments on 11 progressive test scenarios show that HELIOS correctly derives and solves problems from simple rendezvous (8 variables) to multi-leg stay transfers (48 variables), gravity-assist trajectories (17 variables), and solar-sail minimum-time transfers (8 variables). The best compilation success rate reaches 100% (11/11). A multi-model comparison (8 open-source LLM backends, total scores 250-905) verifies the model-agnostic architecture and reveals a positive correlation between model scale and derivation capability.

[AI-244] An Exact Counterexample to Carlsons Associated-Prime Depth Conjecture from a Group of Order 128

链接: https://arxiv.org/abs/2607.23732
作者: Xinan Dai,Wenhao Deng,Yingdong Shi,Tailin Wu,Yuchen Yang
类目: Group Theory (math.GR); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In Question~3.1 of his 1995 paper on depth and transfer, Carlson asked whether the depth of a finite-group cohomology ring is always realized by the dimension of one of its associated primes. We give a negative answer. Let [ G=\SG128859,\qquad k=\kbar. ] An exact presentation certificate proves that \depth H^(G;k)=2 . Okuyama’s associated-prime theorem would convert an associated prime of dimension two into a rank-two elementary abelian subgroup E\leq G satisfying \depth H^(C_G(E);k)=2 . We enumerate all 75 rank-two elementary abelian subgroups of G and obtain six centralizer types. Duflot’s theorem gives depth at least three for four types, while exact ideal-quotient certificates exhibit regular sequences of length three for the remaining two. Hence every rank-two centralizer has cohomological depth at least three, so H^*(G;k) has no associated prime of dimension two. The finite group presentation, the three cohomology-ring presentations, the enumeration summary, and the exact algebraic certificates are included for independent verification. Subjects: Group Theory (math.GR); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.23732 [math.GR] (or arXiv:2607.23732v1 [math.GR] for this version) https://doi.org/10.48550/arXiv.2607.23732 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-245] When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design

链接: https://arxiv.org/abs/2607.23469
作者: Longying Wen,Feiyang Wu,Jinglin Yu,Chongxian Yuan,Renjie Li,Zhaoyu Zhang
类目: Optics (physics.optics); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Applied Physics (physics.app-ph)
备注:

点击查看摘要

Abstract:Photonic-crystal surface-emitting lasers (PCSELs) can combine high-power operation with narrow-divergence surface emission, but optimizing coupled parameters requires costly full-wave simulations. Deep Q-network (DQN) optimization can reuse simulated transitions to guide edits, yet which value-learning mechanisms remain reliable under tight simulation budgets is unknown. We address this gap by comparing baseline DQN and six value-based variants for a seven-variable PCSEL design under a shared objective, simulator, 83-call budget, and four matched initializations. Beyond endpoints, we analyze sample efficiency, policy behavior, and physical response to separate learning gains from favorable starts or exploratory jumps. Dueling DQN is the only variant to improve all four seeds. Relative to the first evaluated designs, its selected structures increase the mean quality factor () from to (), reduce wavelength error by 64%, and increase upward power by 47%; compared with baseline DQN, they achieve a higher mean under the same budget. Other variants yield no consistent improvement; Double DQN reproduces baseline trajectories, while Rainbow-lite shows high upside but strong seed dependence. These results identify Dueling DQN as the most reliable configuration tested for simulation-budget-limited PCSEL inverse design and provide a reproducible framework for attributing algorithmic gains in scientific optimization. The source code is publicly available at this https URL.

[AI-246] A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models UAI2026

链接: https://arxiv.org/abs/2607.23439
作者: Trung Phung,Ilya Shpitser
类目: Methodology (stat.ME); Artificial Intelligence (cs.AI)
备注: 23 pages, 2 figures, accepted for the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)

点击查看摘要

Abstract:Graphical models are ubiquitous in social and empirical science as they are intuitive and easy to use. These models belong to the broader class of Markov models, defined using solely conditional independence (CI) restrictions. In order to estimate finite-dimensional target parameters in such models efficiently, semi-parametric theory provides a principled framework for constructing regular and asymptotically linear estimators via influence functions (IFs). These estimators are asymptotically normal and root- n consistent. Characterizing the class of all influence functions for a target parameter is crucial for statistically efficient inference in these models. For models that are Markov relative to directed acyclic graphs (DAGs), the orthogonal complement of the tangent space is known, implying that for any target the class of all influence functions can be derived once an influence function is obtained. On the other hand, for Markov models not equivalent to a DAG model – such as ordinary Markov models associated with undirected graphs, chain graphs, or acyclic directed mixed graphs – the orthogonal complement has not been characterized, impeding semi-parametric inference in these models. We derive closed form expressions for the orthogonal complement of the tangent space for general Markov models and illustrate our results by characterizing the class of influence functions for the conditional mean parameter in several graphical models. Comments: 23 pages, 2 figures, accepted for the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026) Subjects: Methodology (stat.ME); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.23439 [stat.ME] (or arXiv:2607.23439v1 [stat.ME] for this version) https://doi.org/10.48550/arXiv.2607.23439 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Trung Phung [view email] [v1] Sun, 26 Jul 2026 03:33:14 UTC (51 KB)

[AI-247] FedSLIM: Privacy-Preserving Federated MDL-Based Descriptive Pattern Mining Across Data Silos

链接: https://arxiv.org/abs/2607.23236
作者: Samar Samir Khalil,Noha S. Tawfik,Marco Spruit
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 17 pages, 2 figures

点击查看摘要

Abstract:Federated learning has achieved considerable success for predictive modelling, yet federated descriptive analytics remains largely unexplored. Existing federated pattern mining approaches are predominantly support-based and do not optimise a principled global objective such as Minimum Description Length (MDL). We introduce FedSLIM, the first federated MDL-based framework for descriptive pattern mining. Building on the SLIM principle, FedSLIM enables collaborative optimisation of compact pattern models across distributed databases without sharing raw transactions. We propose two complementary variants that balance privacy, communication, and optimisation fidelity under different deployment assumptions. To evaluate federated MDL mining, we introduce fidelity and discovery-oriented metrics that quantify agreement with a centralised baseline and assess recovery of globally informative patterns. Experiments on multiple real-world datasets under IID and non-IID partitioning show that both variants preserve high-quality compression structure while requiring orders of magnitude less search than the centralised baseline. We further reveal a local-global discovery gap in distributed MDL mining, where globally compressive patterns may be undiscoverable through isolated local optimisation. Both variants recover globally informative patterns absent from all standalone local models, demonstrating the benefits of federated optimisation beyond independent local mining. These results establish federated MDL mining as a practical foundation for privacy-preserving descriptive analytics across distributed data silos.

[AI-248] KAYROS: An Anytime and Exact Open-Source Solver for Duration-Minimization Time-Dependent Vehicle Routing. A Technical Report and a Case Study in Human-AI Engineering

链接: https://arxiv.org/abs/2607.23116
作者: Florian Rascoussier
类目: Optimization and Control (math.OC); Artificial Intelligence (cs.AI); Mathematical Software (cs.MS)
备注: 30 pages, 7 figures, 3 tables. Technical report accompanying the KAYROS solver and Poryos2026 benchmark family

点击查看摘要

Abstract:KAYROS is an open-source solver for duration-minimization time-dependent vehicle routing problems, with or without time windows (TDVRPTW, TDVRP). In these variants, travel times change with departure time, and each route’s dispatch time is a decision. To the best of the author’s knowledge, it is the first openly available solver that is both anytime, streaming improving solutions from the first seconds, and exact, proving optimality with publicly verifiable certificates, for these problems over rich piecewise-linear travel- time functions rather than a time discretization. It has no proprietary dependency and installs with one command. It builds on the state of the art for time-dependent function composition and exact solving, extending the open-source branch-price-and-cut solver of Lera-Romero, Miranda Bront and Soulignac (2020) with an open LP backend, anytime and warm- start behavior, checker-exact pricing, and exact treatment of stepwise travel times. On the MAMUT-routing benchmark collection, KAYROS stands behind 468 published optimality certificates, each requiring agreement among four independent solves, and five certificates strictly improve published reference values. The report also introduces Poryos2026, a benchmark family designed and generated by the author from real OpenStreetMap city road networks. Its 1,080 paired CVRP, VRPTW, TDVRP and TDVRPTW instances combine real road geometries with controlled synthetic demands, time windows and congestion. Every instance carries a checker-validated best-known solution. This report presents the solver, its certification protocol, the benchmark’s generation and feasibility guarantees, and their experimental connection for a broad technical audience. It is also a case study in the intensive human-AI collaboration that made this body of work feasible while keeping its claims independently verifiable.

[AI-249] Exact values and exact upper bounds for families of integers with arithmetic progression intersections (Erdős Problem #272)

链接: https://arxiv.org/abs/2607.23004
作者: Zhanfu Yang
类目: Combinatorics (math.CO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Let t(N) be the largest t for which there exist distinct sets A_1,\dots,A_t \subseteq \1,\dots,N\ such that A_i \cap A_j is a nonempty arithmetic progression for all i \neq j (Erdos Problem #272). Simonovits and Sos proved t(N)=O(N^2) and conjectured \binomN2+1 is best possible; Szabo disproved this by a construction giving t(N) \geq \binomN2+1+\lfloor(N-1)/4\rfloor , proved the asymptotics t(N)=N^2/2+O(N^5/3(\log N)^3) , and asked whether t(N)=\binomN2+O(N) and whether some element lies in all sets of any extremal family (the kernel question). We determine t(N) exactly for all 3 \leq N \leq 12 by exhaustive computation: in this entire range Szabo’s lower bound is exact, and we conjecture that t(N)=\binomN2+1+\lfloor(N-1)/4\rfloor for every N . Towards the matching upper bound we prove, for every N , that Szabo’s bound is the exact maximum over all families with a common element (starred families). The proof combines a self-contained ``defect-one’’ counting inequality for staircase regions with a new structural theorem: every non-progression member of such a family contains a bad pair that no other member can share. Consequently the sharpened conjecture reduces to a single remaining statement, namely Szabo’s kernel conjecture that some element lies in all sets of an extremal family, and we prove first structural constraints on putative non-starred extremal families.

[AI-250] An Explicit Counterexample to Stanleys Rankwise Lower-Bound Conjecture for Differential Posets

链接: https://arxiv.org/abs/2607.22988
作者: Xinan Dai,Yuchen Yang,Wenhao Deng,Yingdong Shi,Tailin Wu
类目: Combinatorics (math.CO); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:In Problem~6 of his 1988 paper on differential posets, Stanley asked for the least possible cardinality of a fixed rank of an r -differential poset and suggested that the minimum should be attained by Y^r , the r -fold Cartesian power of Young’s lattice. We disprove the resulting universal coefficientwise lower bound. For every r\geq 3 , we construct an infinite r -differential poset P^® satisfying [ \cardP^®_4 =\card(Y^r)_4-\left\lfloor\frac r3\right\rfloor. ] For r=3 , the construction replaces thirteen rank-four lower-cover blocks of Y^3 by twelve blocks with the same point and pair incidence multiplicities, producing the initial rank sequence 1,3,9,22,50 instead of 1,3,9,22,51 . A reflection extension then yields an infinite differential poset. The construction does not address the cases r=1 and r=2 . Subjects: Combinatorics (math.CO); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.22988 [math.CO] (or arXiv:2607.22988v1 [math.CO] for this version) https://doi.org/10.48550/arXiv.2607.22988 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[AI-251] Modeling Memory-Dependent Reliability of LLM s: A Hidden Markov Model

链接: https://arxiv.org/abs/2607.22951
作者: Robab Aghazadeh Chakherlou,Siddartha Khastgir,Peter Popov,Xingyu Zhao
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional benchmark-based evaluation, often summarized by aggregate accuracy, provides a point estimate of performance but does not characterize the uncertainty associated with reliability claims. Currently, statistical inference methods for LLM reliability assessment are emerging. However, a key assumption underlying these models is that test outcomes can be treated as independent repeated trials. This assumption may be inappropriate in sequential settings, where later responses depend on earlier interactions through retained context, error propagation, or an evolving interaction state. We extend a hierarchical Bayesian framework for LLM reliability assessment by relaxing the assumption of independent task outcomes and introducing a Hidden Markov Model to capture sequential dependence in benchmark-constructed interaction sessions. In this formulation, outcomes are generated from a latent interaction state evolving according to a first-order Markov process, capturing changes in interaction context. Through experiments using Anthropic Claude and OpenAI on four datasets, we demonstrate the potential impact of sequential dependence on reliability assessment. The results suggest that ignoring sequential dependence may lead to overconfident reliability estimates.

[AI-252] AI-interpreted Optical Scattering for Robust and Focal Depth-Aware Imaging

链接: https://arxiv.org/abs/2607.22867
作者: Eunji Ko,Patrick Ross,Corey Hart,Wolfgang Losert
类目: Optics (physics.optics); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:

点击查看摘要

Abstract:Optical scattering has conventionally been regarded as an impediment in imaging research due to the degradation of image quality during reconstruction. Nevertheless, this study explores two cases in which optical scattering may serve a beneficial role in image reconstruction tasks. We compared the No Scattering MNIST dataset with three Scattering MNIST datasets, each generated under distinct scattering conditions. To assess the information content of the resulting speckle patterns, we employed a Variational Autoencoder (VAE) approach which achieves accuracy comparable to state-of-the-art deep learning approaches, but has an interpretable latent space. We find that scattering can enhance data robustness against spatial pixel loss by effectively distributing information. We also demonstrate that scattering can enable distinctions of focal depth information. We anticipate that these findings will contribute to more efficient imaging techniques, particularly in the presence of obstacles and three-dimensional signals.

[AI-253] Masked Autoencoders Learn Perception-Relevant Representations from Resting State Neural Data AAAI2026

链接: https://arxiv.org/abs/2607.22615
作者: Aleksandr Kovalev,Antonio Lozano,Fabrizio Grani,Cristina Soto Sanchez,Leili Soo,Rocío López-Peco,Adrian Villamarin-Ortiz,Roberto Morollón Ruiz,María del Mar Ayuso Arroyave,Alfonso Rodil,Eduardo Fernández
类目: Neurons and Cognition (q-bio.NC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
备注: 6 papers, 4 figures. NeuroAI Workshop, AAAI 2026

点击查看摘要

Abstract:Clinical neuroprosthetics face a data bottleneck: labeled perception trials are scarce while hours of spontaneous neural activity are largely underutilized. Here, we test whether self-supervised learning can use these unlabeled datasets to improve perception decoding. We pretrained a masked autoencoder on 14.6 hours of spontaneous multiunit activity from an intracortical array in a blind participant’s V1. The model captured interpretable brain structure without supervision: V1’s spatial organization and perceptual state separation both emerged purely from its latent representations. To test these features, we used linear probing (logistic regression on the frozen latents) to measure performance on the data with stimulation. Perception decoding accuracy reached 84.1% on a general psychometric task. On the more difficult threshold-level task, accuracy reached 64.0%. This work shows that spontaneous cortical activity is not noise; it contains rich, task-relevant structure. Unsupervised pretraining on this data is a promising strategy to improve neural decoding.

[AI-254] A Formal Kinetic Theory for Zeroth-Order Newton Dynamics:Stein-Corrected Hessian Estimation and Curvature–Variance Trade-offs

链接: https://arxiv.org/abs/2607.22567
作者: Shihao Ji,Mingyu Li,Zihui Song
类目: Optimization and Control (math.OC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:Zeroth-order Newton-type methods are useful when gradients and Hessians are unavailable, but they behave quite differently from first-order gradient-free methods. We develop a kinetic framework for algorithms that estimate both gradient and Hessian from black-box function values. The naive random-direction Hessian estimator turns out to be biased even on quadratics; a Gaussian–Stein correction is needed to estimate the Hessian of the Gaussian-smoothed objective. Linearizing the inverse Hessian exposes two noise channels: gradient noise preconditioned by the inverse Hessian, and Hessian noise transmitted through an inverse-Hessian sandwich. Under a noisy oracle the second channel carries the second-difference factor \mu_H^-4 . A small-mass kinetic lift links the finite-step Newton update to an underdamped phase-space model; the overdamped spatial limit yields a Lyapunov bound that exposes the curvature–variance trade-off between step size, batch sizes, smoothing radii, and regularization. Numerical experiments confirm estimator identities, the gradient and Hessian variance laws, dimension scaling, inverse-perturbation accuracy, and optimization behavior under query-budget and regularization ablations.

[AI-255] Comparing Optimization Models for Radiotherapy Scheduling

链接: https://arxiv.org/abs/2607.22539
作者: C.C. Rambaldi Migliore,D. Stanicel,N. Musliu,G. Iacca,M. Roveri
类目: Optimization and Control (math.OC); Artificial Intelligence (cs.AI)
备注:

点击查看摘要

Abstract:The Radiotherapy Scheduling Problem (RTSP) involves determining an optimal schedule for patients undergoing radiation treatments, a task that has a massive impact on clinical outcomes given the central role of radiotherapy in cancer care. The daily batch approach–which consists of scheduling all the newly arrived patients together at the end of each day–modelled with Integer Linear Programming, is currently one of the most effective methods for the RTSP. However, this kind of formulation requires substantial computational resources in terms of time and memory. Here, we address these limitations by developing two novel greedy heuristics (named RTSP First Fit and RTSP Best Fit) and use them as constructive heuristics for a Simulated Annealing (SA) approach to optimize the scheduling. The proposed methods–the heuristics alone and their combination with SA–are evaluated on a publicly available dataset against an integer linear program formulation solved with two different state-of-the-art exact solvers. Evaluation metrics include six scheduling objectives capturing patient waiting times, preference satisfaction, and changes in linear accelerator assignment (aggregated in four different weight configurations), solving time, and memory consumption. The results show that the novel heuristics achieve solutions close to those of exact methods, while dramatically reducing runtime and memory usage; furthermore, when combined with SA, they further improve the solution quality while maintaining low runtime and memory usage.

机器学习

[LG-0] Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal Transport

链接: https://arxiv.org/abs/2607.24741
作者: Xinyang Wen
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注: 12 pages, 4 figures, 6 tables

点击查看摘要

Abstract:Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distributed Sinkhorn processes frames sequentially and synchronizes after every iteration. We present TemporalSinkhorn, a parallel-in-time executor that batches future candidates and their repairs without making output accuracy speculative. A centered, row-sharded certificate accepts only a deterministic safe prefix. The remaining candidates share packed Sinkhorn updates; an online projective forgetting rate places audit milestones, while a posteriori residual checks recover from every depth underestimate. Prediction can therefore change work placement but cannot authorize an inaccurate output. On 4 A100 GPUs, a 60-run, five-seed grid at n = 2048 shows that forgetting-guided milestones reduce wall time by 1.15x-1.47x relative to auditing every packed iteration in five statistically resolved regime cells. Against a sequential soft c-transform warm start, temporal execution is 1.42x-3.55x faster across six synthetic streams, with zero marginal-tolerance violations. On Flow Matching minibatch streams, temporal execution is 3.054x-3.632x faster than sequential carry at n = 2048, with no tolerance violations. A separate fixed-kernel test on an RTX 4060 Laptop GPU gives a 4.315x geometric-mean speedup. These are complementary deployment studies rather than a controlled hardware comparison. End-to-end Flow Matching integration, optimized-solver comparisons, and multi-node validation remain open. Comments: 12 pages, 4 figures, 6 tables Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG) Cite as: arXiv:2607.24741 [cs.DC] (or arXiv:2607.24741v1 [cs.DC] for this version) https://doi.org/10.48550/arXiv.2607.24741 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-1] Learning Distributions from Multiple Data Providers

链接: https://arxiv.org/abs/2607.24732
作者: Jon Kleinberg,Amin Saberi,Xizhi Tan,Grigoris Velegkas
类目: Data Structures and Algorithms (cs.DS); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples. The goal is to learn an unknown distribution p on a finite domain [n] . The learner is given a fixed family of queryable sets \mathscrS \subseteq 2^[n] , and each query to S \in \mathscrS returns an independent sample from the conditional distribution p(\cdot \mid S) . Learnability is governed by the co-occurrence graph associated with \mathscrS : two domain elements are adjacent if they appear together in some queryable set. Pointwise consistency is achievable when this graph is connected on the target support. PAC learning requires more: it is possible when the co-occurrence graph is complete. The optimal sample complexity of PAC learning ranges from nearly linear to quadratic. Every query family with complete co-occurrence graph admits sample complexity \widetilde O(n^2/\epsilon^2) , and this bound is tight in the worst case. On the other hand, if [n] is queryable then ordinary sampling improves the bound to \Theta(n/\epsilon^2) , and this cannot be improved further even if every set is queryable. More generally, we identify hierarchical comparabilityas a sufficient structural condition on \mathscr S under which the optimal complexity is nearly linear, \widetilde \Theta(n/\epsilon^2) , with pairwise query families as a canonical example. Finally, the full range of polynomial rates between linear and quadratic is attainable: for every \alpha \in (1,2) , there exists a query family with optimal PAC rate \widetilde \Theta(n^\alpha/\epsilon^2) . Subjects: Data Structures and Algorithms (cs.DS); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG); Machine Learning (stat.ML) Cite as: arXiv:2607.24732 [cs.DS] (or arXiv:2607.24732v1 [cs.DS] for this version) https://doi.org/10.48550/arXiv.2607.24732 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-2] Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs

链接: https://arxiv.org/abs/2607.24726
作者: Justin Sirignano,Konstantinos Spiliopoulos,Samuel Cohen
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:

点击查看摘要

Abstract:The Deep Galerkin Method (DGM) and Physics Informed Neural Networks (PINNs) have become widely-used methods for solving partial differential equations (PDEs) in the rapidly growing field of scientific machine learning. In these methods, a neural network is trained to approximate the PDE solution by using (stochastic) gradient descent to minimize the PDE residual of the neural network. Due to the non-convexity of the PDE residual objective function, the trained neural network may, in principle, only converge to a local minimizer of the objective function (which would not be a solution of the PDE). Therefore, there is a longstanding question regarding the mathematical foundations of these algorithms, and it is highly valuable to establish that the trained neural network will converge to the PDE solution. For a class of semi-linear PDEs (nonlinear in the solution and its first derivative), we prove that neural networks trained with gradient descent to minimize the PDE residual objective function will converge to the PDE solution.

[LG-3] Causal-TS: A Python Library for Causal Discovery in High-Dimensional and Nonstationary Time Series

链接: https://arxiv.org/abs/2607.24673
作者: Mohammad Fesanghary
类目: Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 4 page, Intro paper

点击查看摘要

Abstract:We describe Causal-TS, an open-source Python library for causal discovery in high-dimensional and nonstationary multivariate time series. Causal-TS provides four specialized algorithms-CDNOTS, CDNOTS+, CEDAR, and GRACE-along with wrappers for GES, Granger, LASSO-VAR, and LGES, all sharing a unified conditional independence (CI) test layer with GPU acceleration via PyTorch. A regime discovery pipeline detects structural breaks via pluggable changepoint detectors and runs discovery per regime with regime-specific parameters. A command-line interface, synthetic data generators, and optional DoWhy integration provide an end-to-end pipeline from raw time series to causal effect estimates. The library is pip-installable, tested on Python 3.10–3.12, and available at this https URL.

[LG-4] Explainable Reinforcement Learning via Physics-Aware Policy Distillation

链接: https://arxiv.org/abs/2607.24672
作者: Shaker Al-Tamari,Waled Kadour
类目: Machine Learning (cs.LG)
*备注: 6 pages, 7 figures

点击查看摘要

Abstract:In safety-critical sectors such as robotics and automotive engineering, the deployment of Deep Reinforcement Learning (DRL) is often hindered by the black-box nature of deep neural networks. This lack of transparency poses significant challenges for regulatory compliance and human-agent trust. This paper presents an experimental study aimed at making high-performance continuous control DRL systems interpretable. A policy distillation framework is implemented using the classic Inverted Pendulum benchmark. A high-performance Twin Delayed DDPG (TD3) agent serves as an opaque, continuous teacher model, whose policy is distilled into an interpretable student surrogate based on a shallow Decision Tree. By leveraging a custom physics-aware feature and “Noisy Oracle Rollouts” for dataset generation, the distillation process achieves performance equivalent to the expert teacher. Furthermore, comparative control theory analysis reveals a fundamental trade-off: transitioning from continuous to discrete rule-based control induces high-frequency Bang-Bang actuation and a stable bimodal limit cycle. Simulation results indicate that Bounded-Input Bounded-Output (BIBO) stability is maintained while providing both global and local interpretability for safe autonomous systems.

[LG-5] When Can You Correct Distribution Drift in Temporal Graph Generation? A Sharpening–Drift Tension and an Impossibility for Observation-Based Correction

链接: https://arxiv.org/abs/2607.24662
作者: Tianpeng Li,Xuan Guo,Wenjun Wang,Wang Zhang,Pengfei Jiao
类目: Machine Learning (cs.LG); Social and Information Networks (cs.SI)
*备注:

点击查看摘要

Abstract:Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap. We show this degradation is derivable, general, and not fixable from observations. The masked flow-matching loss decomposes exactly, with no independence assumption, into an irreducible entropy plus a divergence whose derivative along the training path is positive precisely for structures rare during training and common at deployment, diverging as their training probability goes to zero. Empirically the trade-off is a power law with exponent -0.605 ( R^2=0.9977 ), and drift raises the sampler’s error floor without changing how many steps reach it: across seven well-powered conditions the drift-period marginal error varies by at most 6% over a 50\times range of sampling budgets, while the floor sits 2.2\times to 34.3\times above the in-period floor. Because the deployment period is observed, correction looks like a matter of measurement. It is not. We prove that any corrector measurable with respect to past observations leaves at least the conditional variance of the statistic it tracks, and that trend extrapolation beats trusting the last observation only when \mu^2v(1-2\rho) . Both premises are measurable and both go the wrong way: the drift is trendless and mean-reverting, with a one-step innovation as large as the drift itself. An oracle removes 60% of the error, the best observation-based corrector recovers 5.7% of that, and extrapolation is strictly worse than doing nothing clever.

[LG-6] Attribution and Uncertainty Behavior of Learned Residual Gyro Correction for Gyro-Stellar Estimation IJCAI2026

链接: https://arxiv.org/abs/2607.24608
作者: Mariela De Lucas Álvarez,Melvin Laux,Arthur de Freitas Precht,Maurice Martin,Edoardo Caroselli,Frank Kirchner,Alexander Fabisch
类目: Machine Learning (cs.LG)
*备注: 21 pages, 9 Figures, EASi-Explimed Workshop within IJCAI 2026

点击查看摘要

Abstract:This work investigates uncertainty decomposition and explainability in a deep learning-based framework for gyroscope bias correction. A 1-D Convolutional Neural Network is trained to predict residual angular rate corrections from multi-sensor inputs, including gyroscope and star tracker measurements. The bias corrections are sent to a flight-representative Gyro-Stellar Estimator. The network produces both mean corrections and input-dependent (heteroscedastic) aleatoric uncertainty, while epistemic uncertainty is estimated via an ensemble of independently trained models. The proposed approach is trained under nominal conditions and evaluated in both nominal and structured perturbations that include additive and temporally correlated noise. Gradient-based attribution methods are applied to both the correction and uncertainty outputs, enabling a decomposition of the evidence that drives state updates and uncertainty estimates. By aggregating attribution patterns across rotational axes and regimes, we reveal axis-specific behaviors and characterize how structured perturbations influence the collaboration between aleatoric and epistemic uncertainty. Uncertainty analysis shows that aleatoric uncertainty increases with perturbation intensity, but the distributions overlap and the calibration is not consistent across regimes. On the other hand, epistemic uncertainty gives a clear signal that gets clearer as the distributional shift happens, showing that the models disagree more. These results show that aleatoric and epistemic uncertainty work well together and that epistemic uncertainty is better at distinguishing between nominal and perturbed operating conditions. The results provide insight into the behavior of hybrid learning-based state estimation components and motivate the use of uncertainty for downstream monitoring and fault detection. Comments: 21 pages, 9 Figures, EASi-Explimed Workshop within IJCAI 2026 Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.24608 [cs.LG] (or arXiv:2607.24608v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.24608 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-7] PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAM

链接: https://arxiv.org/abs/2607.24583
作者: Kart-Leong Lim
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large scale Bayesian nonparametrics (BNP) learner such as Stochastic Variational Inference (SVI) can handle datasets with large class number and large training size at fractional cost. Like its predecessor, SVI rely on the assumption of conjugate variational posterior to approximate the true posterior. A more challenging problem is to consider large scale learning on non-conjugate posterior. Recent works in this direction are mostly associated with using Monte Carlo methods for approximating the learner. However, these works are usually demonstrated on non-BNP related task and less complex models such as logistic regression, due to higher computational complexity. In order to overcome the issue faced by SVI, we develop a novel approach based on the recently proposed constant stepsize stochastic gradient ascent to allow large scale learning on non-conjugate posterior. Unlike SVI, our new learner does not require closed- form expression for the variational posterior expectatations. Our only requirement is that the variational posterior is differentiable. In order to ensure convergence in stochastic settings, SVI rely on decaying step-sizes to slow its learning. Inspired by SVI and Adam, we propose the novel use of adaptive stepsizes in our method to significantly improve its learning. We show that our proposed methods is compatible with ResNet features when applied to large class number datasets such as MIT67 and SUN397. Finally, we compare our proposed learner with several recent works such as deep clustering algorithms and showed we were able to produce on-par or outperform the state-of-the-art methods in terms of clustering measures.

[LG-8] Evaluating Fuzz Testing for Reinforcement Learning Agents

链接: https://arxiv.org/abs/2607.24577
作者: Zhibin Kang,Hanmo You,Dong Wang,Haiming Zheng,Junjie Chen
类目: Machine Learning (cs.LG); Software Engineering (cs.SE)
*备注:

点击查看摘要

Abstract:Reinforcement Learning (RL) agents are increasingly deployed in safety-critical domains such as robotics, autonomous driving, and drone control, where unexpected behaviors may lead to severe real-world consequences. Fuzz testing has recently emerged as a promising method for exploring the vast state spaces of RL agents and exposing crashes. Although numerous RL fuzzing methods have been proposed, existing studies often differ in evaluation settings, baselines, and metrics, making it difficult to draw reliable conclusions about their relative effectiveness and practical usefulness. To address this gap, we present the first comprehensive empirical study that systematically evaluates RL fuzzing methods from four complementary perspectives: effectiveness, diversity, efficiency, and practical utility. We benchmark five state-of-the-art methods alongside random testing under unified configurations across three environments of increasing complexity (MountainCar, BipedalWalker, and CARLA), and further assess the downstream usefulness of detected crashes for agent robustness improvement and safety monitoring. Our results reveal several key insights. For instance,throughput-oriented methods like MDPFuzz demonstrate superior effectiveness and efficiency in crash discovery, while methods explicitly designed to encourage exploration like SeqDivFuzz excel at uncovering diverse crash behaviors. We also show that fuzzing-generated crashes can meaningfully improve agent robustness and enable accurate safety monitoring with strong cross-method generalization. Beyond these empirical findings, we distill actionable guidance for both researchers and practitioners, highlighting the benefits of combining complementary fuzzing strategies and adopting multi-level diversity analysis to achieve more comprehensive and practical RL testing.

[LG-9] Bit-Accurate FPGA Evaluation of Learned Feature Gating in a Fixed-Point Fourier-Feature Automatic Modulation Classifier

链接: https://arxiv.org/abs/2607.24568
作者: Gawthaman Senthilvelan,Luthira Abeykoon
类目: Machine Learning (cs.LG)
*备注: 8 pages

点击查看摘要

Abstract:Learned feature reweighting can improve automatic modulation classification (AMC) in software, but the same operation introduces additional arithmetic and latency when implemented on an FPGA. This work measures that trade-off in a compact fixed-point classifier using 24 sparse DFT-energy features, 8 phase/statistical features, and a 32-to-128-to-11 multilayer perceptron. A second architecture inserts a learned 32-element, 8-bit, input-dependent gate before the classifier. Gated and ungated models are trained using post-training quantization (PTQ) and quantization-aware training (QAT) with two matched training seeds. The resulting eight checkpoints are compiled independently for an Intel Cyclone V FPGA and evaluated over 352,000 physical-board classifications. Ungated models achieve higher test accuracy in all four matched gate comparisons, with mean gated-minus-ungated differences of -0.784 percentage points under PTQ and -0.616 percentage points under QAT. The effect of QAT changes direction between the two training seeds. In hardware, the gate adds an average of 1,318 adaptive logic modules (ALMs), 1,557 registers, 4 DSP blocks, and 3,140 processing cycles. All 352,000 board predictions agree exactly with an independent integer reference, and 3,760 captured intermediate values from one training seed also match. For this feature representation and implementation, learned gating increases FPGA cost without improving classification accuracy.

[LG-10] he K-SCAN Clustering Algorithm

链接: https://arxiv.org/abs/2607.24537
作者: Filip Kosiorowski,Grzegorz Sroka
类目: Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注:

点击查看摘要

Abstract:In the Big Data era, the scalability of clustering algorithms constitutes a key challenge. Traditional density-based methods (e.g., DBSCAN) offer robustness to noise and the ability to detect non-linear clusters, yet their quadratic time complexity O(N^2) drastically limits their applicability. Conversely, partitional algorithms (e.g., K-Means), with their linear complexity O(N) , impose sphericity on the resulting groups and fail in the presence of outliers. This paper presents K-SCAN – a novel hybrid algorithm that optimizes this trade-off. The method integrates preliminary vector quantization (stochastic Mini-Batch K-Means) to extract a reduced set of weighted micro-clusters, followed by a subsequent density-based structural analysis. Empirical evaluation on datasets of up to 10^6 samples confirms the linear computational complexity of the proposed solution. K-SCAN achieves more than a 3-fold speed-up over the hierarchical BIRCH algorithm, avoiding the costly management of tree-based structures. The method precisely identifies non-linear manifolds while maintaining structural stability (Adjusted Rand Index 0.99), even with noise levels reaching 55% of the data volume. The main limitation of the proposed algorithm, which could not be fully eliminated in the present study, remains its susceptibility to over-smoothing and its difficulty in separating clusters with highly heterogeneous local density. In complex visual spaces, this can lead to the loss of the finest topological details.

[LG-11] From Machine Learning to Large-Scale EO Products: Best Practices for Making Maps ECCV2026

链接: https://arxiv.org/abs/2607.24532
作者: Ghjulia Sialelli,Robin Young,Yuchang Jiang,Cesar Aybar,Linus Scheibenreif,Damien Robert,Clemens Mosig,Adam J. Stewart,Jan D. Wegner,Aleksis Pirinen,Olof Mogren,Konrad Schindler
类目: Machine Learning (cs.LG)
*备注: ECCV 2026 TerraBytes II Workshop paper, non-archival

点击查看摘要

Abstract:Recent years have seen a rapid expansion in the production of large-scale geospatial maps derived from Earth observation (EO) data, driven largely by advances in machine learning (ML) and large computing infrastructure. Although the barrier to generating such maps has dropped substantially, established best practices have yet to emerge, and design decisions made early in the pipeline can quietly propagate errors into the final product. Producing a technically sound and scientifically credible product remains challenging. Choices made at every stage are tightly coupled: preprocessing decisions shape the training signal, dataset design governs what the model can learn and how reliably its performance can be assessed, and global-scale inference introduces engineering challenges in compute and data access at scale, as well as artifact mitigation. Furthermore, uncertainty quantification and independent map validation each require dedicated methodological attention that is often underestimated. This paper presents a concise, end-to-end account of the recommended practices spanning the pipeline from satellite data to an operational map product. We organize the discussion around six interconnected themes: the EO data infrastructure landscape, data selection and preprocessing, ML dataset construction and model training, uncertainty quantification, map production and distribution, and validation. This paper is a condensed version of a longer guide that provides greater depth across all stages, accessible online at this http URL.

[LG-12] Low-Rank Dependence Decomposition via Accelerated Symmetric Non-negative Matrix Factorization

链接: https://arxiv.org/abs/2607.24518
作者: Lavinia Ghita,Dhruv Desai,Jake Goldberg,Roman Yokunda Enzmann
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注: 40 pages

点击查看摘要

Abstract:Symmetric non-negative matrix factorization (SymNMF) recovers latent group structure from a dependence matrix, but its dense, quadratic-memory objective has confined prior work to moderate sizes. We present a large-scale GPU study of seven algorithm families (over 30 configurations) on absolute Pearson correlation and tail pairwise dependence matrices from Extreme Value Theory, two proxies for empirical risk-factor estimation on large portfolios. A trace-identity reformulation eliminates all n \times n intermediates, so a single GPU reaches n \approx 10^5 and multi-node distribution scales to n = 10^6 and beyond. Under a two-phase protocol, eleven methods converge at moderate scale; six remain efficient enough at n = 10^5 (five AdaGrad-family plus ADMM), and five AdaGrad-family methods still converge at n = 10^6 : AdaGrad, RMSprop, and three we introduce (Piecewise AdaGrad, Row-Stochastic SVRG, Block-SVRG AdaptGrow). At n = 10^6 the fastest solver tracks the matrix spectrum: Block-SVRG AdaptGrow wins on the flat, ill-conditioned tail-dependence spectrum, where its lower per-iteration cost decides a long factorization, and full-batch AdaGrad wins on the dominant-low-rank correlation spectrum, where the run is short. We also benchmark spherical K-means as a hard-label baseline: cheaper when angular cluster structure is present, yet provably degenerate once the matrix collapses toward a single common factor, where the soft factorization remains necessary.

[LG-13] Physics Transformer: Tailoring Transformer for General PDE Prediction DATE

链接: https://arxiv.org/abs/2607.24513
作者: Guoze Sun,Rui Zhang,Jiankai Tang,Mengtao Yan,Runze Mao,Zhi X. Chen,Hao Sun
类目: Machine Learning (cs.LG)
*备注: Original version. To be updated

点击查看摘要

Abstract:Transformer architectures have attracted increasing attention for solving partial differential equations (PDEs), owing to their flexibility in handling irregular discretizations and their ability to capture long-range physical dependencies. However, unlike discrete language tokens or fixed-resolution image patches, observed physical fields are finite samples of underlying infinite-dimensional functions. Consequently, effectively applying Transformers to PDEs requires a tokenizer that respects the functional nature of physical fields and constructs physically expressive tokens from arbitrary this http URL this end, we propose \methodnamePhysics Transformer, a function-projection-based Transformer architecture for physical field prediction. Physics Transformer treats a physical field as a continuous function and partitions its discretization into locality-preserving spatial patches. Within each patch, it dynamically learns a set of adaptive local basis functions and projects the sampled field onto these bases to obtain compact physics tokens. The resulting tokens capture diverse latent physical states while preserving fine-scale spatial structures, enabling efficient global interaction through factorized attention across space and physical states. The projected representation further supports efficient decoding at arbitrary query locations. Extensive experiments on diverse benchmarks, ranging from two-dimensional PDE dynamics to industrial-scale three-dimensional CFD simulations, demonstrate that Physics Transformer accurately captures fine-grained physical structures and achieves state-of-the-art predictive performance. These results establish function projection as a practical and effective foundation for designing Transformer architectures for PDE solving.

[LG-14] Context Is King: How In-Context Specification Shapes the Geometry of Concepts

链接: https://arxiv.org/abs/2607.24425
作者: Elad David,Max Fomin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Large language models place structured concepts on geometrically faithful manifolds: weekdays lie on a circle, months on another, usually taken to be a fixed world-model the network stores and looks up. We show that context is king: the structure a model actually uses is set by the in-context specification. A declarative rule fixes not only which relations the geometry encodes but its topology type: the same tokens form a cycle or a branching tree on command, built even on arbitrary, meaning-free tokens with no prior to inherit, which a relabeled stored shape cannot do. When the specification conflicts with a strong pretrained prior, the context-set geometry dominates it in capable models, read from the same activations (representational similarity 0.6–0.9 to the imposed structure versus near-zero to the prior), across the priors we test and both families we study (Gemma, Qwen). Activation patching shows the map is causally used, not a probe correlate: swapping one entity’s activation for another’s makes the model answer with the other entity’s successor under the imposed order. A rough map forms readily, present even in small and base models; what scale gates is using it cleanly: clean dominance and the causal crossover emerge only in the larger models (up to Gemma-31B and Qwen-27B) and weaken or reverse below, so a mechanism present in a large model can be absent in a smaller one of the same family. Whether the model builds this geometry anew or reconfigures a stored one we leave open; operationally, the geometry it uses is the one the context specifies.

[LG-15] K-Survival Means

链接: https://arxiv.org/abs/2607.24405
作者: Abdallah Alabdallah
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In this work, we propose K-SurvMeans, a novel extension of K-Means for clustering survival data. The method explicitly uses the survival outcome in the clustering process to optimize cluster centers, thereby maximizing pairwise survival differences between clusters. The objective function encourages the clusters to be well-separated from the survival perspective. Since the resulting optimization problem is non-differentiable, we employ the Particle Swarm algorithm for the Optimization process. To further improve flexibility and mitigate the curse of dimensionality, we extend the framework to operate in a learned low-dimensional latent space obtained via a dimensionality reduction. This allows the method to capture better-separated clusters and enhance optimization efficiency by reducing the search space. Experiments on multiple publicly available benchmark survival datasets demonstrate that K-SurvMeans consistently yields clusters with improved separation in survival distributions compared to existing deep learning-based survival clustering methods. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.24405 [cs.LG] (or arXiv:2607.24405v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.24405 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-16] When LLM Defenses Backfire: Characterizing Safety Performance and Cost Trade-offs

链接: https://arxiv.org/abs/2607.24392
作者: Tong Zhang,Zexin Li,Simin Chen,Yun Peng
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cost. Rather than treating defenses as a single class, we organize them by operational strategy and examine how different strategies correlate with different side-effect profiles. Across state-of-the-art defense methods, widely used benchmark datasets, and representative open-source LLMs, we find that defenses rarely improve downstream capability, but instead vary in how they trade safety gains against usability and efficiency. In particular, rule-based defenses best preserve task performance, highly conservative self-reflective defenses often increase over-refusal, and multi-round defenses incur the largest runtime overhead. These results provide both a benchmark for evaluating defense side effects and practical guidance for selecting defenses under deployment constraints.

[LG-17] MobiWave: Dispatch-Oriented Graph Wavelets and Drift-Guided Selective Optimization for Autonomous Fleet Rebalancing

链接: https://arxiv.org/abs/2607.24365
作者: Xiao Han,Pinbo Wang,Yuanshao Zhu,Guojiang Shen,Xiangjie Kong
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Autonomous fleets enable mobility platforms to coordinate idle vehicles directly, making fleet-wide rebalancing possible. However, two obstacles limit reliable deployment: overlapping regional and local traffic patterns can hide roads that remain useful for dispatch, and mobility drift can make a trained policy unreliable. Existing spatial aggregation mixes these patterns, while updating all parameters from limited recent data is costly and can damage stable knowledge. We propose \name, a framework that connects a dispatch-oriented multi-scale graph wavelet module with Drift-Guided Layer-Selective Optimization (DGLS). The first module addresses the representation challenge by separating graph-frequency patterns and weighting each scale according to its value for demand prediction and feasible rebalancing. DGLS addresses the adaptation challenge by measuring Dispatch-weighted Spectral Drift, selecting affected layers within a resource budget, and separating short shocks from persistent changes through a drift-aware fast–slow update. Candidate validation further rejects updates that fail to improve held-out dispatch reward without worsening monitored service or safety constraints. Experiments on both real-world datasets and simluated environments demonstrate the effectiveness of \name\ in comparing with state-of-the-art methods. The source code and datasets are available at this https URL.

[LG-18] Perturbative-NeuSA: A Structured Spectral Framework for Time-Dependent PDEs

链接: https://arxiv.org/abs/2607.24345
作者: Xianli Zhu,Jia Yin
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注: 16 pages, 5 figures, 14 tables; supplementary material included

点击查看摘要

Abstract:Neural spectral PDE solvers often learn an entire unresolved vector field even when an inexpensive approximate model can already capture most of the trajectory. Here we introduce Perturbative-NeuSA, a residual formulation that decomposes the target solution into a low-fidelity background and a high-resolution perturbation, so that only the unresolved dynamics is learned. Starting from the exact perturbation equation, the method combines a fixed spectral operator, a background-dependent correction, the background defect in the target PDE, and an optional neural closure. This construction makes the roles of physical structure and neural closure separately measurable. Across 2D Burgers, Klein-Gordon, and heterogeneous 2D wave equations, the deterministic structured solver outperforms the trained NeuSA baseline while requiring no neural-network training. The largest gains occur on Burgers, where the deterministic correction reduces training and extrapolation errors by factors of 24 and 44, respectively. In addition, a Klein-Gordon sweep over seven background resolutions shows that the effect of the closure is conditional: it improves a poor background by 3.6 times, becomes neutral at intermediate resolutions, and degrades a well-resolved background. For the wave equation, however, the closure provides an additional 18% reduction when the remaining residual is interface-localized. Multi-initial-condition diagnostics further show that the useful closure regime depends on the initial-condition spectrum and can disappear in extrapolation when structured correction already captures the dominant Burgers dynamics. Perturbative-NeuSA therefore reframes neural closure as a conditional, diagnosable correction governed by background fidelity, residual organization, and compatibility with the closure model.

[LG-19] Unsupervised Graph Representation Learning with Complementary View Alignment

链接: https://arxiv.org/abs/2607.24338
作者: Zengyi Wo,Shiyu Zhang,Qiyao Peng,Tianpeng Li,Xuan Guo
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Unsupervised graph representation learning aims to derive meaningful node embeddings by capturing both structural and attribute information without relying on labeled data. Existing methods, such as GAEs, have demonstrated effectiveness but typically rely on message-passing mechanisms that assume homophily, leading to performance degradation on heterophilous graphs, where connected nodes exhibit dissimilar features. This homophily bias results in the loss of critical high-frequency components that are essential for identifying heterophilous patterns. To address these challenges, we propose \textscAlignGAE, a novel extension of \textitMaskGAE that preserves the full frequency spectrum through complementary view alignment. Our framework introduces a dual-encoder architecture that separately processes structural and attribute information, incorporates node positional encoding to approximate Neighborhood Identity Distribution (NID), and employs dual reconstruction tasks for both edges and node attributes. We further propose theoretically grounded NID alignment strategies that ensure semantic consistency across views while preserving their distinct characteristics. Through comprehensive spectral analysis, we demonstrate that \textscAlignGAE achieves optimal representation properties when the alignment loss converges. Extensive experiments across 12 benchmark datasets validate our approach, showing that \textscAlignGAE outperforms state-of-the-art methods by up to 18.7% on heterophilous graphs in node classification, while maintaining competitive performance on homophilous graphs. Our results establish a new paradigm for frequency-aware graph representation learning.

[LG-20] DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation

链接: https://arxiv.org/abs/2607.24331
作者: Tan T. Nguyen,Quan V. Dang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually becomes a significant bottleneck as the context window continues to grow. Low-rank compression has recently been studied as an effective approach to reduce KV cache memory while maintaining model performance. However, only a few existing methods treat the Key and Value caches differently, despite their distinct roles. Moreover, these methods typically employ fixed attention-head grouping, which may not fully exploit the structural similarity among attention heads. In this paper, we propose an improved low-rank KV cache compression framework. For the Key cache, we dynamically group attention heads based on Centered Kernel Alignment (CKA) similarity and allocate the rank budget adaptively under a parameter budget. For the Value cache, we adopt the same approach as ReCalKV, refining the low-rank decomposition through offline calibration to improve reconstruction quality. Experimental results on three instruction-tuned LLMs show that our method reduces the number of Key cache parameters while maintaining competitive accuracy. We further observe that the proposed strategy is particularly effective for Multi-Head Attention (MHA) models, whereas it should be applied more conservatively to Grouped-Query Attention (GQA) models, especially in long-context settings.

[LG-21] MEGA-CL: A Molecular Foundation Model for Generalizable ADMET Prediction through Graph External Attention and Contrastive Learning

链接: https://arxiv.org/abs/2607.24314
作者: Tinghui Jin,Kedu Jin,Ying Li,Guanghui Ren,Jingzhi Xue,Shiyu Zhou,Xiaoli Dai,Li-bin Wei,Xijing Chen,Di Zhao,Jinfeng Liu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Predicting the absorption, distribution, metabolism, excretion and toxicity (ADMET) properties of small molecules remains a major challenge in drug discovery. Here, we present MEGA-CL, a foundation graph neural network framework for universal molecular ADMET prediction. MEGA-CL integrates self-supervised contrastive learning with a multi-head external attention mechanism and an enhanced message-passing architecture, enabling simultaneous modeling of local chemical substructures and global inter-graph relationships while mitigating over-smoothing effects commonly observed in deep graph networks. Across 13 benchmark datasets and 21 downstream ADMET tasks, MEGA-CL consistently outperforms state-of-the-art baseline models. In particular, the framework demonstrates robust performance on challenging regression tasks, including clearance (CL) and steady-state volume of distribution (VDss), while maintaining strong generalization ability in independent external validation. Clinically relevant predictive accuracy was achieved, with more than 75% of predictions falling within a 3-fold error range. In an external evaluation on 18 novel compounds derived from recently approved FDA drugs, over 50% of human liver microsome clearance (HLMC) predictions were within a 2-fold error range. To further assess its practical applicability, MEGA-CL was prospectively evaluated on three preclinical drug candidates using in vitro hepatic microsomal metabolism assays and CYP450 inhibition assays guided by model predictions. The predicted HLMC values for all candidates were within 2.5-fold of the experimentally measured values, and 73.3% of CYP450 inhibition endpoints (11/15) were correctly classified. These results demonstrate the potential of MEGA-CL as a generalizable framework for accelerating in silico ADMET evaluation and early-stage drug candidate optimization.

[LG-22] KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems

链接: https://arxiv.org/abs/2607.24260
作者: Shuo Wang,Fang Xi,Wenyuan Huang,Qing Wang,Junming Su
类目: Machine Learning (cs.LG)
*备注: 3 figures, 5 tables

点击查看摘要

Abstract:Modern LLM systems increasingly rely on knowledge-selection processes that produce high-value structured priors, such as ranked evidence, graph topology, multimodal alignment, and confidence signals. Yet LLM serving remains fundamentally oblivious to this rich structure: once such signals are serialized into a prompt, the backend observes only a flat token sequence, forcing dense and uniform consumption of the full key-value (KV) state during decoding. We term this architectural mismatch the Knowledge Selection-Runtime Consumption (KSRC) gap: richer contexts enlarge the full-prompt KV footprint and decode-time memory traffic, increasing latency and degrading throughput even when reasoning depends on only a small fraction of the context. To bridge the gap, we propose Knowledge Access Planning (KAP), a paradigm-shifting execution abstraction that elevates structured knowledge priors from passive prompt-construction hints into first-class physical execution artifacts. KAP establishes a universal intermediate representation (IR)-the runtime access plan-which compiles structured knowledge signals to govern physical KV access without altering logical prompt semantics, model weights, or training procedures. Through this IR, KAP shifts LLM serving from token-aware context consumption to plan-driven, knowledge-aware runtime consumption. We instantiate KAP with GraphSpec, a compiler-executor realization connecting structured knowledge selection to an LLM serving backend. We derive a phase-boundary model for the positive-speedup regime of plan-guided execution. Across 4K-128K long-context QA workloads, GraphSpec maintains answer quality comparable to full-context decoding while decoupling physical KV consumption from prompt length, reducing proposal-time KV access to 5.5% of source KV state at 128K, and fundamentally shifting the scaling trajectory of long-context generation.

[LG-23] Why does Greedy Search produce Optimal Clustering Outcomes? A Fixed-Core Assignment Theory

链接: https://arxiv.org/abs/2607.24237
作者: Kai Ming Ting,Kaifeng Zhang,Sanjay Chawla
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Many existing clustering methods are designed based on a set-oriented definition—a cluster is a set of similar points—relying a point-to-point similarity function to find similar points. This works well for compact clusters, but clustering performance can deteriorate badly when cluster shapes are irregular, and densities or sizes vary between clusters. Recent `Cluster-as-Distribution’ (CaD) clustering has been shown to discover these generic types of clusters in practice by treating each cluster as a set of independent and identically distributed points generated from some unknown distribution via a greedy search, achieving a clustering objective equivalent to that of Spectral Clustering, but with better clustering outcomes without eigen-decomposition. However, a theoretical analysis of this phenomenon is still lacking. Our analyses are from two angles. First, we analyze the approximation error between the true and empirical distribution embeddings. Second, we show that the greedy search employed to achieve the CaD clustering objective can be mapped to a partition matroid—yielding greedy optimality. These yield a near-optimality guarantee for the CaD clustering objective, with regret controlled by the approximation error. This is the first analysis that explains why CaD clustering via greedy search can discover clusters of arbitrary shapes, densities and sizes (where all set-oriented clustering methods have failed to discover) when the estimated cluster embeddings faithfully approximate the underlying cluster distributions.

[LG-24] EXE-Bench: Ranking the Tradeoffs of AI-based Windows Malware Detectors for Real-World Usability

链接: https://arxiv.org/abs/2607.24177
作者: Andrea Ponte,Daniel Gibert,Matous Kozak,Dmitrijs Trizna,Maura Pintor,Battista Biggio,Fabio Roli,Luca Demetrio
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Due to the lack of systematic evaluations, we are not yet able to determine which AI-based Windows malware detector to deploy in production, since existing evaluations (i) differ in terms of data used for both training and testing; (ii) do not consider temporal analysis to showcase whether models withstand the passage of time; (iii) avoid security evaluations with adversarial attacks that could highlight their brittleness against content-injection attacks; and (iv) neglect the computational requirements for deployment, risking slow inference on endpoints. For these reasons, we develop EXE-Bench, a comprehensive benchmark of AI-based Windows malware detectors. EXE-Bench assesses performance, temporal and adversarial robustness, and computational overhead, aggregating them into a single score for direct and fair model comparison. Through EXE-Bench, we highlight how evaluations conducted only after deployment are suboptimal and unable to provide a complete picture of their performance. In particular, through our analysis, we remark how much domain knowledge instilled through feature engineering is still extremely useful in this domain, resisting both time and adversarial attacks, in stark contrast with most of the deep networks that only excel right after deployment.

[LG-25] Forecasting the Emergence and Evolution of Crash Hotspots: A Unified Deep Learning Framework for Proactive Traffic Safety

链接: https://arxiv.org/abs/2607.24168
作者: Jingwen Zhu,Keshu Wu,Pei Li,Steven T. Parker,Bin Ran,David A. Noyce
类目: Machine Learning (cs.LG); Emerging Technologies (cs.ET)
*备注:

点击查看摘要

Abstract:Road crashes remain among the gravest threats to public safety, and preventing them is a defining task of transportation systems worldwide. Much of that harm concentrates at hotspots, yet a hotspot is less a place than an episode; it emerges quietly at an intersection or along an arterial, intensifies for weeks, then subsides, only to reappear elsewhere. Enforcement guided by maps of past crashes inevitably trails this cycle, patrolling yesterday’s hotspots while tomorrow’s form unwatched. Breaking that lag requires three capabilities at once: detecting hotspots as they are born, forecasting where they will sit next week, and following each one through its life. We introduce HERALD (Hotspot Emergence, Risk Anticipation, and Life-cycle Dynamics), a unified deep learning framework that provides all three from a single statewide model. HERALD distills each county’s recent crash history into weekly risk maps and forecasts the next with a CNN–Transformer, whose mixture-of-experts lets one model serve dense urban cores and sparse rural corridors alike. Each forecast is anchored in the county’s long-run crash geography, sharpened by the self-exciting effect of recent crashes, and paired with explicit warnings of where new hotspots are about to appear. Followed over time, every hotspot acquires a legible life story, from birth through growth and stability to decline and death. Across six heterogeneous Wisconsin counties, HERALD forecasts more accurately than five identically trained baselines, locates hotspots most precisely, and flags emerging risks before they take hold. A single adjustable setting trades accuracy for extra sensitivity where deployment demands it. The result shifts hotspot management from mapping the past to anticipating the future.

[LG-26] Calibrated Tree-Neural Fusion for Fine-Grained Vegetation Community Classification

链接: https://arxiv.org/abs/2607.24160
作者: Dristi Datta,Md Khalid Hasan Sakib,Manoranjan Paul
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Accurate vegetation-community classification is essential for ecological monitoring, habitat assessment, and evidence-based environmental management in heterogeneous landscapes. Existing studies often rely on standalone tree ensembles or generic neural networks, although fine-grained ecological classes frequently exhibit overlapping spectral, topographic, and structural characteristics. Many frameworks also provide limited protection against stacking leakage, insufficient probability calibration, weak minority-class evaluation, and little evidence of stability across repeated data splits. To address these limitations, this study proposes Calibrated EcoTreeFuseNet-Plus, a tree-neural probability-fusion framework that combines out-of-fold tree probabilities, EcoFuseNet-V2 outputs, validation-selected meta-learning, and post-hoc temperature scaling. Raster values from six LiDAR-derived terrain and canopy variables and two hyperspectral vegetation indices were extracted at coordinate-based reference locations. Quality control removed 26 samples with missing elevation and one sample with non-finite NDWI, producing 1,833 complete records across 29 vegetation and non-vegetation classes. On the held-out test set, the proposed model achieved an accuracy of 0.8000, a macro F1-score of 0.7768, a balanced accuracy of 0.7903, and an MCC of 0.7903. Calibration reduced the expected calibration error from 0.3866 to 0.0651 without changing class predictions. Five-seed evaluation yielded a macro F1-score of 0.7717 +/- 0.0112, indicating stable performance across repeated splits. The results demonstrate a reliable discrimination-calibration trade-off for small-sample, fine-grained ecological classification.

[LG-27] An Empirical Study of Feature Selection Granularity

链接: https://arxiv.org/abs/2607.24145
作者: Muhammad Rajabinasab,Arthur Zimek
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Feature selection aims to identify the most informative and relevant features for a given dataset, either in terms of capturing the underlying data structure and distribution better, or with respect to the performance on a downstream task. Existing research in this area has largely focused on developing novel algorithms (in both supervised and unsupervised settings), proposing new evaluation metrics and frameworks, or benchmarking the performance of existing methods. In this work, we examine feature selection through an algorithmic design perspective. Conventional feature selection algorithms typically compute feature importance scores globally across the entire feature set and then select the top-ranked features in a single step. However, this approach raises a critical question: Can the presence of less informative (or noisy) features mask or obscure the true importance of other, more relevant features? In other words, would a recursive strategy, where features are removed one by one while re-evaluating importance at each step, yield different and potentially better results than the standard global ranking approach? To answer this question, we conduct an extensive empirical study using five diverse feature selection algorithms. We implement each algorithm under both the conventional global selection design and the greedy recursive elimination design. We then analyze the impact of this algorithmic choice, both individually for each method and collectively across all methods, on a range of standard feature selection evaluation metrics. The empirical evaluation results show that the greedy approach improves the overall feature selection quality almost consistently, albeit on the expense of higher computational cost, supporting our initial expectation that the curse of dimensionality also obscures the ways of mitigating it.

[LG-28] MAPLE: Efficient and Diverse Multi-Alpha Generation for Portfolio Construction

链接: https://arxiv.org/abs/2607.24131
作者: Yu-Chen Den,Kuan-Yu Chen,Kendro Vincent,Tien-Hao Chang
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注:

点击查看摘要

Abstract:Classical alpha mining achieves strong risk-adjusted returns by combining many low-correlated predictive signals, yet deep learning stock-ranking methods typically produce a single alpha per stock, rely on increasingly complex architectures with diminishing gains, and obtain diversity only through separate models or implicit routing, without explicitly controlling inter-alpha correlation. We introduce MAPLE (Multi-Alpha Position-aware Listwise Ensembling), a backbone-agnostic framework that recovers this diversity principle within a single training pass. MAPLE combines a unified, capacity-scaled prediction head with an extreme-rank weighted listwise ranking loss and a diversity regularizer that explicitly penalizes pairwise correlation across alphas. Across four equity markets spanning the US, China, and Japan, MAPLE achieves the best average Sharpe and Calmar ratios among nine baselines, using up to 55x fewer parameters and 2.5x less training time, and generalizes across five backbone architectures with Sharpe and Calmar Ratio gains of 10-23% and 17-43%, respectively. Behavioral analysis further shows why each component works: the unified head already reduces inter-alpha correlation before any diversity loss is applied, and the extreme-rank loss lets diversity regularization improve rather than erode per-alpha ranking quality as capacity scaling sustains this balance at scale. These results show that principled loss design and capacity allocation, rather than architectural complexity, drive diverse and effective multi-alpha generation.

[LG-29] EmBed-T: A Multi-Dimensional Benchmark for Table-Level Embeddings VLDB2026

链接: https://arxiv.org/abs/2607.24130
作者: Ayeen Poostforoushan,Liane Vogel,Carsten Binnig
类目: Databases (cs.DB); Machine Learning (cs.LG)
*备注: Accepted to the 4th International Workshop on Tabular Data Analysis (TaDA) @ VLDB 2026

点击查看摘要

Abstract:Tabular data is the dominant structured-data modality, and learning table representations has become a core research direction. Table-level embeddings in particular underpin a wide range of applications, including table retrieval, data lake discovery, and table classification. Despite their importance, there is still limited understanding of how different embedding approaches behave across tasks, making systematic evaluation and analysis essential. In this work, we introduce a systematic evaluation of table-level embeddings that captures several complementary properties required for downstream effectiveness. We realize this evaluation by extending TEmBed, a recently proposed testbed for tabular embeddings, whose table-level coverage is currently limited to a single retrieval task. An empirical study over the TEmBed model pool confirms that no single model excels across all tasks, demonstrating that table-level embedding quality cannot be reduced to retrieval alone.

[LG-30] Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation

链接: https://arxiv.org/abs/2607.24083
作者: Valerio Belli(UNIROMA, UCL),Valerio Modugno(UCL),Enrico Mingo Hoffman(HUCEBOT),Fabio Amadio(HUCEBOT)
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注:

点击查看摘要

Abstract:Reinforcement learning can produce robust humanoid controllers, but each new task is typically trained as a separate policy with its own reward design and training process. Motion imitation provides an alternative source of motor competence by training policies to track retargeted human motions, yet the resulting controllers remain reference trackers and are not directly usable as task policies. We propose a three-stage pipeline that turns motion-imitation skills into a reusable hybrid motion prior (HMP) for humanoid locomotion. First, an expert policy is trained to imitate retargeted human motion-capture clips. Second, the expert is distilled into a frozen architecture composed of a proprioceptive encoder, a residual vector-quantized (RVQ) codebook, and an action decoder. Third, task-level policies are trained to solve locomotion tasks by selecting discrete codebook entries while the HMP remains frozen. We evaluate the method on velocity tracking, point-goal navigation, and fall-recovery velocity tracking in simulation, and deploy the velocity-tracking policy on a real Unitree G1 robot. The distillation process preserves the tracking behavior of the expert, while the resulting HMP can be reused without retraining as the action interface for different downstream locomotion policies. The learned HMP reveals an interpretable codebook structure in which the number of active RVQ stages modulates the available gait patterns. We further show that training the codebook with the rotation trick improves latent organization and reduces downstream falls compared with a standard straight-through estimator.

[LG-31] Constrained Reinforcement Learning Using Successor Representations

链接: https://arxiv.org/abs/2607.24057
作者: Michael Girstl,Alexander Mattick,Christopher Mutschler
类目: Machine Learning (cs.LG)
*备注: published in Transactions for Machine Learning Research 2026

点击查看摘要

Abstract:Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to introduce an additional cost signal in the Markov Decision Process, which notifies the agent of unwanted behavior independently of the reward signal. Unfortunately, current methods are hard to adapt to changes in the cost function introduced by, e.g., domain shift or obstacles moving over time. The lack of adaptability means that policies are too unflexible to deal with complex real-world conditions. We propose the Safe Deep Successor Representation (SafeDSR), a novel method that allows quick retraining of policies towards new cost structures. SafeDSR extends the Deep Successor Representation (Kulkarni et al., 2016) to Constrained Reinforcement Learning by introducing a single learnable weight matrix to decouple the learned value function across dynamics, rewards, and costs. This matrix can be updated in a supervised manner instead of having to adapt the whole network if the cost structure of the environment changes. We demonstrate this ability in a freely configurable two-dimensional navigation environment and show that our method is competitive on a simple navigation task while being considerably more flexible

[LG-32] Beyond Local Inspection: Global Guideline-Grounded Evaluation of Post-hoc XAI Methods for ECG Classification

链接: https://arxiv.org/abs/2607.24035
作者: Nils Gumpfer,Michael Guckert,Samuel Sossalla,Birgit Aßmus,Jennifer Hannig
类目: Computers and Society (cs.CY); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Explainable AI (XAI) is used to assess whether artificial intelligence models rely on meaningful patterns, yet explanations that appear plausible for individual predictions may systematically misrepresent model behavior. This is particularly problematic in medicine, where models may rely on irrelevant signal characteristics rather than disease-specific patterns without being recognizable. We address this challenge using electrocardiogram (ECG) data, for which clinical guidelines provide explicit knowledge about diagnostically relevant signal regions. We introduce a global, guideline-grounded framework that aggregates explanations across heartbeats to evaluate them against clinically defined regions of interest. Using four binary classifiers trained on PTB-XL, we assess 13 gradient-based methods across two categories of patterns: low-amplitude segments and high-amplitude QRS morphology. Our results reveal a systematic failure of methods transferred from computer vision. Their explanations often follow signal amplitude rather than clinical relevance, with mean Spearman correlations up to 0.69, leading them to overlook diagnostically decisive low-amplitude regions. For ischemia, LRP- \epsilon assigns only 4.6% of relevance to the ST segment, compared with 63.8% for LRP-SIGN. Nine of 13 methods fall below chance for at least one condition, indicating inconsistent reliability across patterns. These findings show that global, domain-grounded evaluation can uncover systematic explanation failures not obvious from sample-level heatmaps.

[LG-33] When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility Calibration and Cost KDD KDD2026

链接: https://arxiv.org/abs/2607.24010
作者: Pin Qian,Su Wang,Chong Peng,Junxian You,Lifei Liu,Haoran Yu,Yihang Chen,Xiaochong Jiang
类目: Machine Learning (cs.LG)
*备注: Accepted at the ACM SIGKDD KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI; 7 pages, 1 figure, and 4 tables

点击查看摘要

Abstract:Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive retrieval. Yet evaluations often leave the operating point underspecified: two systems may both claim a 50% evidence-usage budget while realizing different held-out usage rates, so higher accuracy can reflect a looser budget rather than a better retrieval policy. We study budget-aware evaluation for Active RAG by recasting active retrieval as utility estimation, where retrieval is valuable only through its marginal correctness change over a no-retrieval answer. This view separates three questions that single-point evaluations conflate: whether trigger scores rank useful retrieval decisions, whether thresholds calibrated on past data meet future budgets, and how trigger-side computation changes deployment cost. We operationalize these questions with exact top-k utility frontiers, deployable threshold frontiers, conservative budget frontiers, harm audits, and cost decompositions. Across knowledge-intensive multi-hop QA datasets and open instruction models, retrieval harm is non-negligible, router rankings change across datasets and budgets, nominal thresholds can miss target usage, and simple uncertainty or retrieval-score baselines often rival learned utility routers. Budget-aware Active RAG evaluations should therefore report frontiers, realized usage, threshold-transfer error, harm rates, and cost decompositions alongside accuracy.

[LG-34] Disentangling Acoustic Cues in Alzheimers Pathology and Perception: The Roles of Language and Gender INTERSPEECH2026

链接: https://arxiv.org/abs/2607.23977
作者: Liu He,Yuanchao Li,Yin-Long Liu,Rui Feng,Yiming Wang,Jiaxin Chen,Yizhe Wang,Jiahong Yuan
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注: Accepted at Interspeech 2026

点击查看摘要

Abstract:Acoustic biomarkers show promise for detecting Alzheimer’s Disease (AD), yet whether the cues driving diagnostic AI align with those salient to human listeners is underexplored across languages and genders, where pathological markers and perceptual strategies differ. We train models to predict clinical AD status (pathology) and human perceptual scores across Mandarin and Greek, male and female speakers. Using SHAP for interpretability and statistical models for validation, we compare feature importance by subgroup. Results reveal a context-dependent divergence: pathological-perceptual alignment is significant for Mandarin and female speakers but disappears for Greek and male speakers, where pathology models did not exceed chance; this is a failure mode that population-specific auditing surfaces. Global Explainable AI (XAI) explanations can mask critical demographic divergences, highlighting the need for population-specific explainability auditing for equitable deployment of clinical speech AI.

[LG-35] Joint Flow Matching for Generator-Consistent Classification

链接: https://arxiv.org/abs/2607.23946
作者: Hayden McAlister,Lech Szymanski
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We introduce Joint Flow Matching (JFM), a training framework for continuous normalising flows over multiple variables. Standard flow matching transports variables from noise to data simultaneously, offering no natural mechanism for forward and reverse conditional inference from a shared joint model. JFM resolves this by assigning opposite roles to each variable at the temporal endpoints. We prove that JFM produces a consistent joint distribution where that forward or reverse integration are conditionals of the same joint. We explore this consistency in the context of joint classification and generation as the basis for interpretability in discriminative-generative models. We validate JFM on conditional datasets producing competitive accuracy with inherently well-calibrated confidence scores without post-hoc calibration, and classifier-consistent image generation.

[LG-36] Variational Boosting for Physics-Informed Neural Networks

链接: https://arxiv.org/abs/2607.23940
作者: Pavlos Protopapas,Kaylee Vo
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Physics-Informed Neural Networks (PINNs) solve differential equations by minimizing the residual of a nonlinear operator over a neural parameterization of the solution. However, monolithic PINNs often suffer from ill-conditioning, spectral bias, and optimization instability. We introduce a variational boosting framework in which solutions are constructed additively in function space. Each stage trains a weak learner whose converged correction satisfies a local orthogonality condition, equivalent to a projected functional gradient descent step onto the tangent space of the network’s function manifold. Because each correction network is deliberately small, the restricted minimization admits full Newton or conjugate gradient updates, which are typically infeasible in large PINNs. The resulting method separates global nonlinear refinement into a sequence of well-conditioned subproblems while preserving the full variational structure of the operator. This framework provides a geometric interpretation of multi-stage PINNs as projected functional gradient descent and enables stable second-order optimization for nonlinear differential equations. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.23940 [cs.LG] (or arXiv:2607.23940v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.23940 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-37] DECAF: De-Clustering for Adaptive Representational Unlearning

链接: https://arxiv.org/abs/2607.23934
作者: Anjie Le,Can Peng,Hongcheng Guo,J. Alison Noble
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Machine unlearning, which aims to remove the influence of specific training data from a trained model, is a key requirement for privacy, accountability, and adaptive deployment. We argue that many unlearning methods are vulnerable to a simple clustering attack, which can recover class structure in an unsupervised manner, limiting their suitability for continual deployment where removal requests must be handled reliably on demand. To address this, we propose DECAF (DE-Clustering for Adaptive Forgetting), a post-hoc method that operates only on the forget set and is designed to break the cluster. DECAF combines input noise, confidence suppression, and entropy-based output diversification to disrupt the residual feature-space structure associated with forgotten data. On CIFAR-10 with ResNet-18, DECAF attains 0.10% forget-class accuracy, 79.4% retain accuracy, and an AUS of 0.88, surpassing all other baselines. In cluster-based analysis, it attains performance comparable to that of unlearning methods that use the full training set, while being significantly more efficient. Code: this https URL.

[LG-38] Greedy dynamical meta-learning

链接: https://arxiv.org/abs/2607.23925
作者: Aria Yom
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Gradient descent scales well to large models, but becomes unstable over long time horizons. Gradient-free optimizers can scale to arbitrary timespans, but are hobbled by high dimensions. Since learning occurs in large models over long timescales, neither of these approaches is likely to produce traits which can accelerate the learning process. Instead, we propose a meta-learning algorithm in which the agent learns to modify its own weights and biases. Our algorithm consists of an inner loop, wherein the agent performs some high-dimensional optimization upon itself, and an outer loop, wherein we perform some low-dimensional optimization upon the inner loop. Since the outer loop handles very few parameters, standard zeroth-order methods may be used.

[LG-39] WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

链接: https://arxiv.org/abs/2607.23909
作者: Sen Wang,R. Gnana Praveen,Bidhan Roy,Marcos Villagra
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注: 9 pages, 4 figures

点击查看摘要

Abstract:Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four LIBERO simulation suites, WorldDiT lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites. These results provide a strong sub-billion-parameter baseline for future scaling studies.

[LG-40] ADVERSARIAL: And-Inverter Graph-Assisted Hardware Trojan Detection At Scale

链接: https://arxiv.org/abs/2607.23882
作者: Yaroslav Popryho,Debjit Pal,Inna Partin-Vaisband
类目: Machine Learning (cs.LG); Hardware Architecture (cs.AR); Cryptography and Security (cs.CR)
*备注: 9 pages, 5 figures, 3 tables. Accepted for publication at the IEEE/ACM International Conference on Computer-Aided Design (ICCAD 2026)

点击查看摘要

Abstract:Modern System-on-Chip (SoCs) often contain hundreds of millions to tens of billions of gates, making existing Hardware Trojan (HT) detection methods impractical due to their immense scale. The proposed approach incorporates symbolically enabled learning by modeling flattened gate-level netlists as Boolean networks represented as And-Inverter Graphs (AIGs), where all internal nodes are 2-input AND gates and inversions reside on the edges. Each directed connection is expressed as a triple within a Knowledge Graph Embedding (KGE) framework, producing compact, constant-size per-node representations that retain multi-hop structural context. The AIG’s bounded fan-in and uniform semantics ensure training and inference complexity scale linearly with edge count, addressing major scalability bottlenecks in HT detection. Symbolically enabled learning across deep datapaths enables the model to differentiate circuit structures from rare and functionally inconsistent connections that signify potential Trojan triggers and payloads. Experiments on large-scale SoC benchmarks demonstrate clear geometric separation between Trojan and benign nodes and practical scalability.

[LG-41] Flash-CNNCap: Capacitance Extraction via Image Mapping

链接: https://arxiv.org/abs/2607.23877
作者: Hector R. Rodriguez,Jiechen Huang,Wenjian Yu
类目: Machine Learning (cs.LG)
*备注: Accepted to ICCAD-26

点击查看摘要

Abstract:We present Flash-CNNCap, a CNN-based capacitance extractor that reformulates full-matrix capacitance prediction as image-to-image regression over spatial contribution maps. Prior scalar CNN-based extractors require O(n^2) forward passes to recover all pairwise capacitances in a window with n conductors. Flash-CNNCap replaces the scalar target with dense contribution maps: a total-capacitance model and a master-conditioned coupling model each predict a spatial map that is reduced to conductor-level values through mask aggregation, cutting full-matrix reconstruction to O(n) passes. The resulting totals and symmetrized pairwise couplings define the corresponding Maxwell-style capacitance matrix under the standard off-diagonal sign convention. The maps are learned from conductor-level labels without per-pixel supervision. An ablation study over 13 model configurations selects a U-Net that matches ResNet baselines on total capacitance (1.5-3.1% MARE) and achieves the strongest coupling accuracy (3.0-4.6% MARE) across all evaluated CapBench subsets, with a 17.5\times full-matrix speedup on windows containing 134 conductors on average. A deployed pipeline reads Design Exchange Format (DEF) geometry and writes Standard Parasitic Exchange Format (SPEF) output, processing 1,024 windows in 51.23 seconds with a 4.4\times speedup over OpenRCX on the same benchmark. Code and trained models are available at this https URL.

[LG-42] XMix: Combating Extremely Noisy Labels via Local Smoothness in Self-Supervised Feature Space

链接: https://arxiv.org/abs/2607.23865
作者: Chengqi Li,Yangdi Lu,Zhihao Shi,Wenbo He,Chamseddine Talhi,Nadjia Kara
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Supervised deep learning models rely on large, accurately labeled datasets, yet noisy annotations are often unavoidable and can severely degrade performance under high noise levels. Recent state-of-the-art methods tackle this by using sample selection strategies that exploit the memorization effect to filter out clean data for semi-supervised learning. However, these methods struggle with extreme noise, class imbalance, and require careful tuning or prior noise knowledge. To address these limitations, we propose XMix, a novel framework that leverages local smoothness in the self-supervised feature space to systematically enhance all stages of the sample selection process, without dependence on potentially corrupted labels. First, XMix estimates the noise rate using maximum likelihood among self-supervised feature neighbors. Second, these neighbors then help identify additional clean samples and ensure balanced selection across classes during sample selection. Finally, in the semi-supervised learning phase, XMix uses neighboring samples to generate more reliable pseudo-labels. Our empirical results show that XMix substantially outperforms existing methods in extremely noisy environments and maintains superior performance in standard LNL benchmarks.

[LG-43] Covariance Last-Layer Ensembles: Function-Space Diversity for Efficient Uncertainty Quantification

链接: https://arxiv.org/abs/2607.23856
作者: H. Martin Gillis,Isaac Xu,Gabriel Spadon,Thomas Trappenberg
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:A Last-Layer Ensemble (LLE), K linear units on one shared frozen feature map, is an efficient single-pass approach to the disagreement-based epistemic uncertainty for out-of-distribution (OOD) detection. Its weakness is that members share the backbone gradient and can converge toward the same function, collapsing the inter-member diversity the signal depends on. Whether last-layer diversity can be restored, and what mitigates the collapse, is an open question. The weight-orthonormality defining Orthonormal Certificates (OC), the weight-orthonormal special case of the LLE, is only an indirect correction; it decorrelates the weights of the members, not their predictions. Here, we instead target the collapse directly in function space, with a Covariance Last-Layer Ensemble (cov-LLE) that places a direct covariance penalty on member activations. Cov-LLE restores the function-space diversity that weight-orthonormality cannot, and at matched K recovers much of the diversity and calibration of a deep ensemble at 1\times backbone cost (in-distribution prediction variance 0.05!\to!9.3 vs.\ 22.1 ( \times10^-3 ), and ECE 0.135!\to!0.090 vs.\ 0.035 , for a K\times -cost deep ensemble), at no cost to accuracy. Viewing OC as a last-layer ensemble also organizes detectors into a two-axis taxonomy (by how their units are trained and how their outputs are scored) and exposes the OC score as a magnitude, motivating a scale-invariant, label-free direction score that repairs its near-OOD failure, adding +0.16 to +0.18 ROC AUC on every backbone.

[LG-44] DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

链接: https://arxiv.org/abs/2607.23822
作者: Yuhang Wang,Lingyao Li,Hao Zhou
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specific regularities for driver-specific style. We introduce DriveDNA, a large-scale naturalistic dataset and benchmark for personalized driving-style modeling, comprising 4,121 drives from 465 drivers across 115 vehicle models and totaling 975 hours of human-controlled driving at 10 Hz with forward video, collected from community drivers in everyday use. DriveDNA defines driving style as a consistent, driver-specific behavioral pattern in how a vehicle moves under similar conditions. The benchmark evaluates this signal through three core tasks: few-shot driver re-identification, personalized behavior prediction, and condition-matched comparison, and provides behavioral annotations plus 276,248 rule-generated maneuver events across six classes with large-scale human auditing. We evaluate baselines spanning classical descriptors, supervised and self-supervised time-series encoders, multimodal fusion, probabilistic prediction, and zero-shot foundation models under a fixed multi-seed protocol. Learned representations substantially outperform classical descriptors on unseen drivers (AUROC .935 vs. .707) and retain driver-specific information under matched driving conditions, while descriptor performance approaches chance. Video-only models achieve comparable re-identification accuracy but exhibit severe route leakage, showing that strong recognition may arise from contextual shortcuts rather than driving behavior. These findings show that reliable driving-style evaluation must assess both the behavioral value of learned representations and their robustness to vehicle, drive, and condition confounds.

[LG-45] SCTA: An Agent ic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing KDD KDD2026

链接: https://arxiv.org/abs/2607.23821
作者: Shuyu Chen,Chen Zhu,Ye Zhang,Yang Li,Qiqi Xie,Haohan Wang
类目: Machine Learning (cs.LG); Genomics (q-bio.GN)
*备注: 10 pages, 4 figures. To appear in the ACM SIGKDD Workshop on Data Mining in Bioinformatics (BioKDD 2026)

点击查看摘要

Abstract:Identifying therapeutic target genes from single-cell RNA sequencing (scRNA-seq) data remains a fundamental challenge in translational biology. Unlike bulk assays, scRNA-seq captures heterogeneous cellular states and rare subpopulations, but this same heterogeneity makes target discovery highly sensitive to analytical choices throughout the pipeline, including preprocessing, cell population selection, differential expression analysis, and downstream biological interpretation. As a result, existing workflows and general-purpose analysis agents often produce unstable or difficult-to-interpret target hypotheses, limiting their reliability for disease-focused discovery. We present SCTA (Single-Cell Target Agent), a decision-centric agentic framework for stable and interpretable target gene discovery from scRNA-seq data. Rather than treating analysis as a single general-purpose reasoning task, SCTA decomposes target discovery into specialized agents aligned with key decision points in the single-cell pipeline and constrains downstream reasoning with structured biological evidence. In a representative ablation study on hereditary chronic pancreatitis, we demonstrate that SCTA’s full evidence integration yields the most stable target selection across independent runs among the tested configurations, while recovering biologically coherent, disease-relevant mechanisms validated in prior studies. These results suggest that decision-aware agent orchestration tailored to the structure of single-cell analysis can improve the robustness, interpretability, and practical utility of target discovery in precision medicine.

[LG-46] GNN-based Multi-Agent Control of Traffic Shockwaves in Sparse Vehicular Ad-hoc Networks

链接: https://arxiv.org/abs/2607.23792
作者: Prachi Nandi,Madhuri Malakar,Sonakshi Satpathy,Pabitra Mohan Khilar
类目: Networking and Internet Architecture (cs.NI); Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:

点击查看摘要

Abstract:Traffic shockwaves are stop-and-go waves that propagate upstream through the streams of vehicles and are one of the major causes of traffic congestion, fuel inefficiency, and increased accident rates in modern transportation systems. Although Connected and Autonomous Vehicles (CAVs) offer a promising opportunity to mitigate such shockwaves, most existing control strategies rely on global traffic state information, making them impractical for early-stage deployment of Vehicular Ad-hoc Networks (VANETs). In this paper, we propose a decentralized Multi-Agent Reinforcement Learning (MARL) framework that integrates a Graph Neural Network (GNN) to enhance the control architecture of connected and autonomous vehicles. The proposed approach enables vehicles to learn cooperative control policies using locally available information and interaction with neighboring vehicles. The effectiveness of the proposed scheme is evaluated using a scalable simulation environment under realistic highway traffic conditions. Simulation results show that the proposed GNN-based MARL framework can reduce the propagation of traffic shockwaves by up to 80%, even when only 10% of the vehicles are connected.

[LG-47] A Comparison of Data Augmentation Methods for Training Deep Neural Networks on Synthetic Aperture Sonar

链接: https://arxiv.org/abs/2607.23770
作者: C.J. Moore,Gregory D. Vetaw,Jordan Malof
类目: Machine Learning (cs.LG)
*备注: This paper was originally published in the proceedings of the International Conference on Underwater Acoustics 2026

点击查看摘要

Abstract:In this work we study Automatic Target Recognition (ATR) for Synthetic Aperture Sonar (SAS) data with a focus on deep neural networks (DNNs). The main challenge in training DNNs for SAS ATR arises from the limited quantity of labeled target examples, which arises due to the significant costs and time required to collect real-world SAS data. One successful general strategy for mitigating the problem of limited training data is augmentation, which generates additional synthetic training data by introducing realistic variations to available real-world data. Prior research has investigated a variety of augmentation strategies for SAS ATR, including conventional image augmentations (e.g., contrast changes, cropping) as well as physics-based augmentations, which are motivated the specific physics of SAS data. Building on prior work, we systematically compare many of these existing augmentation strategies for training DNNs for SAS ATR. We also investigate the impact of augmentation when combined with modern DNN architectures such as transformers. The results indicate that augmentation can improve target recognition accuracy, although benefits vary, and not all augmentations are beneficial.

[LG-48] On the post-hoc Evaluation of PDE Discovery: A Multifaceted Challenge of Scientific Advancement

链接: https://arxiv.org/abs/2607.23753
作者: Baptiste Mathevon,Farah Cherfaoui,Amaury Habrard,Marc Sebban
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Partial differential equation (PDE) discovery aims to identify from data the governing law of a physical system. Constituting a cornerstone of scientific advancement, it has become during the past decade a major line of research in the rapidly evolving field of Physics-informed Machine Learning (PiML). Among the remaining open problems to address in this domain, the post-hoc evaluation of discovered PDEs raises the particular difficulty of being multifaceted. Indeed, it requires jointly considering predictive accuracy, physical consistency, interpretability, and out-of-distribution generalization capacity. Given that some of these properties are conflicting, it is worth noting that the wide range of existing evaluation metrics only partially address the overall problem, potentially leading to overly interpreted conclusions about the validity of a presumed new physical theory. From an abundant literature spanning machine learning, numerical analysis, information theory or symbolic regression, we propose, to our knowledge, the first taxonomy of PDE evaluation metrics, and discuss their advantages and limitations in depth. Based on the observation that evaluation is often achieved on a case-by-case basis and that a universally accepted methodology remains elusive, we further provide recommendations with the aim of promoting standardized and reliable practices, before sketching promising future lines of research in this field. We argue that this paper is intended both for ML experts who design new PDE discovery algorithms and for users of these methods aiming, in real applications, to discover and validate well-founded scientific laws.

[LG-49] Soft-Constrained Optimization of Latent Space in Variational Autoencoders

链接: https://arxiv.org/abs/2607.23751
作者: Ye Shi
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注: 12 figures, 7 tables

点击查看摘要

Abstract:The usefulness of a variational autoencoder (VAE) depends on two properties of its latent space that are hard to obtain together: high encoding capacity in the individual latent variables, and a low-dimensional, disentangled organization of those variables. Weakening the Kullback-Leibler regularization raises capacity but degrades disentanglement, while strengthening it prunes latent variables away entirely. We formulate VAE training as a soft-constrained optimization problem that addresses both. First, we impose an entropy-based constraint (EC) on individual latent variables, showing that the entropy of a latent code upper-bounds the mutual information it carries about the generative factors of the data. Second, we propose a weight-filter method that exploits the slack of the soft constraint to prune low-entropy dimensions during downstream training. On dSprites, the EC raises the aggregate latent-variable activation score by 43-62% over a vanilla VAE, attains the highest FactorVAE score among the \beta \beta-VAE variants (0.891 vs 0.847), and lowers reconstruction error by up to 38%. On MNIST, the weight filter reduces the latent dimensionality supplied to a downstream classifier from ten to two while holding accuracy above 90%, converging in 37% fewer epochs than the same procedure without the EC. We also find that low-entropy discrete factors tend to merge into a single latent variable, whereas high-entropy continuous factors are distributed across several.

[LG-50] RUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

链接: https://arxiv.org/abs/2607.23734
作者: Muhammad Umar Farooq Qaisar,Lin Zhang,Zhen Chen,Wajdy Othman,Shehzad Ashraf Chaudhry,Chang Liu
类目: Networking and Internet Architecture (cs.NI); Machine Learning (cs.LG)
*备注: 7 pages, 3 figures, submitted to IEEE Internet of Things Magazine

点击查看摘要

Abstract:Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.

[LG-51] Outcome-Confounded Local Supervision in On-Policy Distillation

链接: https://arxiv.org/abs/2607.23731
作者: Guoqing Ma
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:On-policy distillation (OPD) trains a student on its own trajectories while a teacher supplies dense token-level likelihoods at student-visited prefixes. These likelihoods are often read locally: agreement appears safe to imitate, whereas disagreement appears to identify an error. We show that both readings are confounded by the outcome of the completed trajectory. We introduce an outcome-resolved diagnostic that crosses pointwise teacher-student divergence with final-answer correctness, separating safe imitation, productive divergence, harmful divergence, and agreement-on-failure. In an eight-seed mathematical-reasoning study with a Qwen3-8B student and Qwen3-32B teacher, agreement-on-failure constitutes 67.84% of pooled response-token mass; with a Qwen2.5-7B/32B pair it remains 67.68%. The result persists across threshold, sequence-level, format, and truncation audits. Even on prompts that the Qwen3 teacher solves in all four independent attempts, student accuracy rises to 86.91% but agreement-on-failure remains 14.76%. We then run three matched training probes that use the available signals to imitate, mask, or contrast whole trajectories; none consistently reduces agreement-on-failure. The result points to a localization limitation: local divergence paired with a trajectory-level outcome does not identify where a failed trajectory became unrecoverable. Addressing this limitation requires additional positional information, such as process labels, teacher continuations from student prefixes, or token-level alignment across rollouts. Our contribution is therefore diagnostic rather than a new training method.

[LG-52] Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning

链接: https://arxiv.org/abs/2607.23726
作者: Zahra Abdalla Elashaal,Afef Hfaiedh,Nahla Khraief,Issmail Ellabib,Giansalvo Cirrincione
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 27 pages, 6 figures, 11 Tables

点击查看摘要

Abstract:Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning. To address these challenges, we propose a two-level Hierarchical Reinforcement Learning (HRL) framework. The first level handles high-level strategic planning, while the low-level uses the continuous-control Soft Actor-Critic (SAC) algorithm, and they utilize entropy-regularized policy optimization. The proposed framework was trained and evaluated using the Search-and-Rescue-2 (SAR-2) dataset. HRL-SAC effectively addresses sparse-reward long-horizon search problems characterized by delayed rewards and continuous control, and its outperforming the flat SAC baseline reinforcement learning in terms of success rates, coverage efficiency, and convergence. These findings indicate that hierarchical entropy-regularized policies are a promising solution to tackle long-horizon sparse-reward reinforcement learning tasks.

[LG-53] he Intruder Threshold: A Spectral Law for LoRA Fine-Tuning

链接: https://arxiv.org/abs/2607.23711
作者: Peng Xie
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:LoRA fine-tuning can create intruder dimensions: new leading singular vectors of the updated weight matrix W+BA that are nearly orthogonal to all pretrained singular vectors and that drive catastrophic forgetting. Since their discovery, no theory has predicted, layer by layer on measured spectra, when they appear. We derive a per-layer critical update strength s^\ast=\bar\theta/(\gamma\sigma_1(BA)) , computed from the measured spectrum of W alone through the rectangular spiked-deformation transform, together with an exact secular-equation characterization of the updated spectrum, with no fitted parameters. In a pre-specified study spanning four dense Transformer families, a state-space model, a mixture-of-experts model, and an encoder-decoder (18 adapters, 9,840 layer scans), the law localizes the empirical threshold within a factor of two on 82% of layers, separates intruder-bearing from intruder-free layers at deployment with a mean AUC of 0.89 , holds unchanged on six third-party adapters, and predicts where WikiText-2 perplexity begins to degrade; a combination of the two pre-specified edge evaluations reaches 98% and is confirmed out-of-bag on the external adapters ( 0.997 ). Full fine-tuning disperses its update far below the threshold of every layer, which resolves the asymmetry between LoRA and full fine-tuning. Norm-matched interventions confirm that threshold-crossing layers, rather than update magnitude, carry the forgetting, and a spike-budget rule derived from the thresholds, requiring one SVD and no validation sweeps, reduces forgetting by 62% on the most fragile model at no task cost.

[LG-54] Extreme Volatility Warning under Label Scarcity via Multi-Source Anomaly Fusion

链接: https://arxiv.org/abs/2607.23682
作者: Jin Qian,Zhangzhi Xiong,Mingrui Li,Zhen Liu
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Early warning of extreme market volatility is central to financial risk management, but actionable events are rare, nonstationary, and often triggered by exogenous information shocks. In our CSI~300 setting, only \sim 80 positive samples are observed across 791 training days, making heavily supervised multi-source models unstable. We first analyze a 100K-parameter hierarchical text-signal fusion model (HTSF) and find that added parameterization hurts in this low-label regime. Motivated by this failure, we propose \textbfAAMSF (Anomaly-Augmented Multi-Signal Fusion), a semisupervised framework that combines Isolation Forest anomaly scores over market indicators, GDELT events, Chinese financial news, and English media with lightweight Ridge score fusion. We further introduce \textbfT-AAMSF, a temporal extension for multi-day anomaly accumulation. On CSI~300 (2018–2023), AAMSF achieves test AUC-ROC \textbf0.680, outperforming the strongest unsupervised baseline (0.630) and neural baseline (0.588), while T-AAMSF improves PR-AUC to 0.291. Ablations reveal strong source asymmetry: GDELT and domestic financial news provide complementary risk signals, whereas English media consistently reduces performance, and learned weighting is unreliable under validation noise. These results suggest an empirical design principle for label-scarce financial risk warning: robust anomaly geometry and source reliability can matter more than supervised representation capacity.

[LG-55] Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set

链接: https://arxiv.org/abs/2607.23679
作者: Heyang Zhao,Tianyuan Jin,Weixin Wang,Vincent Y. F. Tan,Pan Xu,Quanquan Gu
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Recent years have witnessed increasing interests in tackling heteroscedastic noise in bandits and reinforcement learning. In these works, the cumulative variance of the noise \Lambda = \sum_t=1^T \sigma_t^2 , where \sigma_t^2 is the variance of the noise at round t , is used to characterize the statistical complexity of the problem, yielding \emphsimple regret bounds of order \tilde\calO(d \sqrt\Lambda / T^2) for d -dimensional linear bandits with heteroscedastic noise. However, with a closer look, \Lambda remains the same order even if the noise is close to zero at half of the rounds, which indicates that the \Lambda -dependence is not optimal. In this paper, we revisit the stochastic linear bandit problem with heteroscedastic noise, where the action set is prefixed throughout the learning process. We propose a novel variance-adaptive algorithm \textttVAEE (Variance-Aware Exploration with Elimination) for large action set, which actively explores actions that maximizes the information gain among a candidate set of actions that are not eliminated. With the active-exploration strategy, we show that \textttVAEE achieves a \emphsimple regret with a nearly \emphharmonic-mean dependent rate. For finitely many actions, we propose a variance-aware variant of G-optimal design based exploration, which achieves a simple regret with sharper dependence on d . We also establish a nearly matching lower bound for the fixed action set setting indicating that \emphharmonic-mean dependent rate is unavoidable. To the best of our knowledge, this is the first work that breaks the \sqrt\Lambda barrier for stochastic linear bandits with heteroscedastic noise.

[LG-56] No Free Lunch in Flow Surrogates under Time-Varying Boundary Conditions: A Two-Regime Study

链接: https://arxiv.org/abs/2607.23667
作者: Georg Winkler,Martin Stoll
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Fluid Dynamics (physics.flu-dyn)
*备注: 19 pages

点击查看摘要

Abstract:A flow surrogate validated on a simple regime is often taken as evidence that the approach will carry to a richer one. We test this assumption on two transient flows under time-varying boundary conditions emulating the process startup: the three-dimensional slurry film in chemical-mechanical planarisation (CMP), a core semiconductor-manufacturing process, and the two-dimensional Karman vortex street (KVS) behind a cylinder. Eight surrogate models are compared on one shared evaluation pipeline, differing in whether they learn the full field or a latent representation, and whether they predict trajectories in one shot or step by step. No single architecture wins both regimes. On the film, a one-shot full-field model reconstructs the process-relevant cumulative wall shear stress to 3.2% relative error. On the wake, a latent autoregressive DeepONet retains 96% of the shedding power that direct and one-shot models damp to almost zero. The deciding axis is the treatment of time. The self-sustained wake requires the phase memory that autoregressive feedback provides, while the boundary-driven film rewards a direct map. Pointwise RMSE picks the wrong model in both regimes, so the evaluation scores five physical questions instead, the field, its structure, invented motion, amplitude, and timing. The trained surrogates answer queries 10^3 to 10^4 times faster than the finite-element solver, but the offline cost of the training simulations means they pay off from the first query beyond the training set for CMP and the third for the KVS. The choice of surrogate should follow the dynamical character of the target flow, and its validation should use failure-mode-resolved metrics, since neither the winning architecture nor its validation transfers.

[LG-57] DP-IVON-Gradsq: Differentially Private Squared-Gradient Improved Variational Online Newton ECML KDD2026

链接: https://arxiv.org/abs/2607.23649
作者: Nour Jamoussi,Ikram Dridi,Giuseppe Serra,Marios Kountouris
类目: Machine Learning (cs.LG)
*备注: Accepted at the 2nd Workshop on Machine Unlearning and Privacy Preservation (WIPE-OUT 2), held in conjunction with ECML PKDD 2026

点击查看摘要

Abstract:Differential privacy provides formal privacy guarantees for training neural networks on sensitive data, while Bayesian deep learning offers a principled framework for uncertainty-aware prediction. Combining these two objectives remains challenging, as privacy noise can interact with the stochasticity introduced by Bayesian posterior sampling. In this work, we investigate differentially private variational Bayesian learning through the Improved Variational Online Newton (IVON) optimizer. We introduce DP-IVON-Gradsq, a private variant of IVON. The proposed method constructs its curvature estimate from the privatized gradient using a noise-corrected squared-gradient estimator, reducing the direct interaction between posterior-sampling noise and privacy noise while preserving the Adam-like computational efficiency of IVON. We evaluate DP-IVON-Gradsq on CIFAR-10 against the standard private optimizers DP-SGD and DP-Adam over a range of privacy budgets. The results show that DP-IVON-Gradsq is competitive under weak-to-moderate privacy constraints, i.e., large-to-moderate values of \varepsilon , while degrading under strong privacy. Code is available at this https URL.

[LG-58] Optimal Reward Shaping: Autonomous Car Parking Case Study

链接: https://arxiv.org/abs/2607.23617
作者: Emre Özkaya,Nicolas R. Gauger
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注: 12 pages, 8 figures. Includes supplementary video demonstration and open-source code link

点击查看摘要

Abstract:Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this work, we present a parameterized reward shaping framework featuring coverage-gated alignment feedback, drive-direction switch regularization, and an aligned episode termination mechanism evaluated on an autonomous parallel parking task. Crucially, we show that environmental reward parameters and algorithmic hyperparameters are deeply co-dependent, requiring joint meta-optimization to achieve stable convergence. By employing surrogate-based Bayesian optimization, our co-optimized Deep Q-Network (DQN) agent resolves characteristic control failure modes, significantly outperforming uncalibrated baselines across both success rate and trajectory smoothness.

[LG-59] Restoration Flow Matching-Based Channel Refinement and Equalization Correction for MIMO Semantic Communications

链接: https://arxiv.org/abs/2607.23615
作者: Wenkai Liu,Nan Ma,Jianqiao Chen,Xiaodong Xu,Meixia Tao,Ping Zhang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:In multiple-input multiple-output (MIMO) semantic communication, imperfect channel state information (CSI) and equalization mismatch can seriously degrade semantic reconstruction quality. To address this issue, we propose a unified restoration flow matching (RFM)-based framework for channel refinement and equalization correction. Specifically, the channel RFM (CRFM) module is developed to refine the coarse channel, thereby improving channel estimation accuracy. Based on the refined channel, the developed semantic RFM (SRFM) module is employed to correct the residual distortions in the post-equalization latent space. The key idea is to formulate the two cascaded inverse problems of channel estimation and equalization as the unified conditional restoration task, in which the learned conditional velocity field guides the perturbed distribution towards the target distribution. To enhance the robustness of these two modules under various distortion conditions, we develop a dual-anchor perturbation training strategy that jointly learns near-manifold refinement and large-error correction, and implement inference through a few-step deterministic ordinary differential equation (ODE) solver. Extensive experiments on MIMO channels and visual semantic transmission tasks demonstrate that the proposed scheme improves key metrics for channel estimation and semantic reconstruction quality. Moreover, compared with representative diffusion-based generative baselines, the proposed method requires fewer sampling steps.

[LG-60] Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter IROS2026

链接: https://arxiv.org/abs/2607.23565
作者: Yuchao Mei,Guohao Zhang,Luxia Ai,Haopeng Chen,Wenbing Tao
类目: Robotics (cs.RO); Machine Learning (cs.LG)
*备注: 8 pages, 7 figures. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

点击查看摘要

Abstract:Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision. To address this, we propose an anticipatory risk-guided reinforcement learning framework. Leveraging privileged simulator states, we construct a directionally aligned future collision risk map based on the Closest Point of Approach (CPA). Through an asymmetric actor-critic architecture, the network is trained to self-predict this structured risk, which explicitly guides the visual policy during deployment. A lightweight spatio-temporal encoder extracts motion cues directly from onboard depth sequences, bypassing explicit object tracking or optical flow estimation. Extensive simulated and real-world experiments demonstrate that our method effectively improves safety margins and flight efficiency in dense dynamic clutters compared to existing baselines. Furthermore, the learned policy achieves robust zero-shot Sim-to-Real transfer on a physical quadrotor, relying purely on abstracted spatio-temporal depth sequences and its self-predicted risk priors, validating the effectiveness of our approach and its robust generalization from simulation to reality.

[LG-61] Random Forest-Based Prediction of Bone Volume Fraction and Fracture Position from S-Parameters

链接: https://arxiv.org/abs/2607.23563
作者: Jianhe Li,Jinsui Meng,Yida Zhao,Zihe Wang,Liaoran Sun,Tao Shan
类目: Machine Learning (cs.LG)
*备注: 5 pages,6 figures,This is a summary report prepared by undergraduate students I supervised, formatted as a letter

点击查看摘要

Abstract:In this paper, we propose a method for predicting bone volume fraction (BVF) and fracture position by constructing a random forest model based on multichannel S-parameters. A nine-antenna microwave scanning system is designed and fabricated to acquire the multichannel S-parameter data. Bone-mimicking phantoms are developed, and corresponding experiments are conducted to validate the effectiveness of the proposed approach. Both synthetic and experimental results demonstrate the validity of the method.

[LG-62] Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling

链接: https://arxiv.org/abs/2607.23518
作者: Hengyuan Cao,Shizhuo Cheng,Mingxuan Liu,Weicheng Huang,Yunhong Lu,Chenxi Cai,Yan Zhang,Min Zhang
类目: Machine Learning (cs.LG); Biomolecules (q-bio.BM)
*备注:

点击查看摘要

Abstract:The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination. However, existing approaches are predominantly implemented under a single-target, single-state assumption, limiting their ability to model multi-target or multi-state interactions required for advanced function-oriented protein design. Here, we introduce Chamaileon, which unifies multi-target and multi-state binder design by formulating the problem as cross-context binding landscape modeling. The framework is underpinned by a training paradigm termed In-Context Complex Co-Design (I3CD) for context-aware sequence-structure co-modeling. During inference, we employ Mixture-of-Paths Sampling (MoPS), a scalable strategy that optimizes a single sequence across contexts while alleviating the scarcity of high-quality multi-conformational paired data. Extensive evaluation on our newly constructed benchmark, CROSS, demonstrates that Chamaileon effectively generates sequences adaptable to diverse conformational landscapes and multi-target requirements. The code is available on this https URL.

[LG-63] opological Data Analysis and Graph-Theoretic Approaches for Tennis Match Prediction

链接: https://arxiv.org/abs/2607.23509
作者: Jake Schwaderer,Alexander Bastien,Omid Khormali,Alejandro Navarrete,Mia Pesavento,Angelika Elderbrook
类目: Machine Learning (cs.LG); Algebraic Topology (math.AT); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:We present two approaches for predicting tennis match outcomes using topological data analysis and graph theory on ATP singles matches from 2000-2025. The first method applies lower-star filtration to player competitive networks, extracting topological features through persistent homology using four summary methods (VAB, HNAV, HWNAV, OW-HNPV) combined with Modified Band Depth analysis. Algorithmic optimizations including ego graph approximations and triangle elimination enable analysis of about 66k matches. Our Random Forest model achieves 66.2% accuracy (AUC = 0.719) using topological, graph-theoretic, and ranking features. Feature importance analysis reveals that rankings contribute 36.3%, centralities 25.5%, and TDA features 24.0%, with topological features providing complementary signal. When rankings are unavailable, the topology-only model maintains 63.56% accuracy, demonstrating that network-derived features alone capture meaningful competitive structure. The second method uses a modified Katz similarity index with temporal edge weighting, achieving 62.48% accuracy on held-out test data. This work represents the first application of lower-star filtration to tennis prediction, provides systematic comparison of four topological summary methods in sports analytics, and demonstrates that TDA can achieve above-chance prediction using network topology alone while providing additional value when combined with traditional features.

[LG-64] Physics-Informed Neural Networks for Discovering Periodic Orbits in the Gravitational Three-Body Problem

链接: https://arxiv.org/abs/2607.23501
作者: Nikolaos Kollias,Nikolaos Matzakos
类目: Machine Learning (cs.LG); Chaotic Dynamics (nlin.CD); Computational Physics (physics.comp-ph)
*备注: 39 pages, 16 figures, 16 tables

点击查看摘要

Abstract:Locating periodic solutions of chaotic dynamical systems normally requires an initial guess close enough to the target orbit for numerical continuation or gradient-based search to converge. We show that Physics-Informed Neural Networks (PINNs) trained on sparse, noisy observations \emphwithout initial conditions recover periodic orbits of the gravitational three-body problem, including orbit families absent from the training data. The method rests on a second-order ODE formulation, fixed-frequency Fourier features, percentile-based adaptive refinement, and a trainable scaling parameter, each validated on forward problems. Across two 100-seed ensembles, 23 – 25% of runs converge to families not present in the training data. We then ask what determines which family emerges. Two \chi^2 tests give a consistent answer: changing the training data source significantly shifts the distribution of recovered families ( p 0.001 , Cramér’s V = 0.339 ), whereas switching between the two initialization distributions tested does not ( p = 0.620 , V = 0.094 ). The random seed selects which family a given run recovers; the \emphdistribution the weights are drawn from does not shift the aggregate frequencies, but the training data does. The evidence is empirical: we do not characterize the loss landscape analytically, and PINNs remain slower than conventional integrators on well-posed initial-value problems. What the experiments establish is that the recovered orbits are verifiable rather than merely plausible: the identified ones refine to genuine periodic solutions, a network trained on Lagrange data recovers the figure-eight choreography (Li–Liao class I.A.1, matched to seven significant digits in T^* ), and one trained on figure-eight data recovers a Broucke–Hadjidemetriou–Hénon orbit closing to \delta_T 10^-9 .

[LG-65] An adaptive multi-fuzzy logic model for diagnosing transformer faults using dynamic weight optimization

链接: https://arxiv.org/abs/2607.23486
作者: Kim-Anh Nguyen,Huy Hoang Le,Ba Tu Phung
类目: Machine Learning (cs.LG)
*备注: This paper has been published in e-Prime - Advances in Electrical Engineering, Electronics and Energy. Please cite the published version

点击查看摘要

Abstract:Dissolved gas analysis (DGA) is crucial for diagnosing early power transformer failures. Traditional DGA interpretation methods like Duval Triangle, IEC ratio, Roger ratio, Doernenburg ratio and Key Gas are inconsistent and vary in accuracy, especially for multiple fault conditions. We propose an Adaptive Multi-Fuzzy Logic (AMFL) model integrating multiple DGA methods with fuzzy logic and a dynamic weight adjustment mechanism. Unlike existing approaches with fixed weights, this system iteratively evaluates each method’s diagnostic performance, identifies multiple fault types, and adjusts weights based on fault prediction accuracy. A feedback-based optimization recalibrates weights after each cycle to ensure optimal solution convergence. The model, implemented in MATLAB/Simulink, is validated against DGA datasets with known error conditions. Results show the AMFL model significantly improves diagnostic accuracy, especially in complex error scenarios, and enhances adaptability to new datasets. Comparative analysis demonstrates the proposed method outperforms traditional fixed weight multi-fuzzy systems in accuracy, consistency, and reliability of error detection. This work provides a robust, flexible diagnostic tool for transformer condition monitoring and supports more accurate asset management decisions.

[LG-66] Charging Phase Health Indicators for Battery State-of-Health Estimation: A Systematic Comparison of CC CV and Combined Approaches under Cross-Battery Validation

链接: https://arxiv.org/abs/2607.23482
作者: Huy Hoang Le,Kim-Anh Nguyen
类目: Machine Learning (cs.LG)
*备注: This paper has been published in Eksploatacja i Niezawodnosc. Please cite the published version

点击查看摘要

Abstract:Accurate State-of-Health estimation is essential for safe battery operation and cost-effective maintenance. Although numerous health indicators have been derived from constant-current (CC) and constant-voltage (CV) charging phases, their effectiveness under realistic cross-battery validation remains insufficiently studied. This work addresses this gap through a systematic comparison of CC-only, CV-only, and combined indicator sets using rigorous Leave-One-Battery-Out (LOBO) validation on the NASA battery aging dataset. Four CV-phase indicators and CC phase duration are evaluated individually and in combination. Results show that the combined CC+CV approach achieves the best performance (R2 = 0.874), confirming that CC and CV phases capture complementary degradation information. Moreover, a 119% performance gap is observed between standard 5-fold cross-validation and LOBO validation, indicating that conventional evaluation overestimates practical accuracy. Based on these findings, practical guidelines are provided for indicator selection under data and computational constraints.

[LG-67] A Multi-stage Constrained Optimization Framework for Data-driven Problems

链接: https://arxiv.org/abs/2607.23480
作者: Ye Shi
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注: 14 pages, 10 figures, 6 tables, 3 algorithms. Preprint

点击查看摘要

Abstract:Variational autoencoders (VAEs) transform high-dimensional, often noisy data into a compact latent representation, making downstream optimization more tractable. Three challenges persist in VAE-based constrained optimization: (i) sampling effectively within the latent space, (ii) identifying the active decision variables that actually influence the objective and constraints, and (iii) enforcing constraints without destabilizing training. We propose a Multi-stage Constrained Optimization Framework (MCOF). First, an entropy-constrained VAE (EC-VAE) coupled with a feature selector embeds objective and constraint information into a designated subset of latent variables, so that optimization proceeds over a low-dimensional subspace while the remaining coordinates supply solution diversity. Second, a Uniform Transformation (UT) module applies a per-dimension probability integral transform, replacing the irregular aggregate posterior with a uniform distribution over a bounded box and mitigating posterior collapse and Gaussian mixture bias. Third, a constraint-priority filter method (CPFM) solves the resulting surrogate problem by alternating violation-reduction and objective-reduction steps under a filter acceptance test, returning solutions that are feasible for the learned surrogate to a specified tolerance without requiring multiplier estimation. Finally, unselected latent coordinates are resampled to generate diverse decodings of a single optimized solution. We validate MCOF on a synthetic problem, where we ablate each stage and recover the analytic optimum, and on a ZINC250k drug design task, where the generated molecules satisfy the imposed constraints and are entirely novel relative to the training set.

[LG-68] Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning

链接: https://arxiv.org/abs/2607.23474
作者: Minh Vu,Konstantinos Slavakis
类目: Machine Learning (cs.LG)
*备注: This work has been submitted to the Elsevier for possible publication

点击查看摘要

Abstract:This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions (S-GMM-QFs). The framework reconciles streaming, non-stationary data with the Riemannian structure of the parameter space while handling distributional mismatch through experience replay. S-GMM-QFs are introduced via Hadamard overparametrization, enabling interpretable sparsification through smooth regularization that facilitates Riemannian-based optimization. Overparametrization allows the framework to adaptively identify meaningful components from a large initial pool, yielding sparse models where interpretability emerges naturally from geometry: each component’s parameters (means and covariances) explicitly encode its geometric role in the ambient state-action space. These geometric roles are learned through online gradient descent on a smooth objective over a (Cartesian-product) Riemannian manifold. Numerical tests demonstrate that S-GMM-QFs match or exceed deep RL methods while using substantially fewer parameters and achieving faster improvement per observed transition. Notably, parameter efficiency and interpretability combine to maintain strong generalization in low-parameter regimes where sparsified deep RL approaches degrade.

[LG-69] Learning to Optimize: Joint Routing and Flow Allocation on Sparse Non-Euclidean Networks

链接: https://arxiv.org/abs/2607.23467
作者: Haomiao Sun,Fang He,Congyuan Ji,Xindi Tang
类目: Machine Learning (cs.LG)
*备注: 34 pages, 14 figures, and 10 tables

点击查看摘要

Abstract:We study an integrated pickup-and-delivery problem on sparse, non-Euclidean networks that jointly optimizes cyclic routing, cargo flow allocation, and cross-cycle service. The tight coupling of these operational constraints creates a complex discrete-continuous decision space with highly restricted feasible regions. To overcome these computational challenges, we propose Double-Channel Graph Attention (DCGA), an end-to-end reinforcement learning framework. DCGA isolates network reachability and demand-service logic into separate graph channels and constructs valid routes using a simulator-coupled, constraint-informed decoder. Experiments on LinerLib benchmarks demonstrate that DCGA achieves seconds-level inference and delivers state-of-the-art solution quality on instances beyond a specific scale, with its advantage over existing baselines widening significantly as problem size increases. Supported by extensive stability and ablation analyses, our results demonstrate that this structure-aware learning approach provides an effective, low-latency engine for realistic routing-and-flow optimization.

[LG-70] Extending Fourier Neural Operators for Modeling Parameterized and Coupled PDEs ICLR2026

链接: https://arxiv.org/abs/2607.23466
作者: Cheng Jing,Uvini Balasuriya Mudiyanselage,Abhishek Verma,Kallol Bera,Shahid Rauf,Kookjin Lee
类目: Machine Learning (cs.LG)
*备注: Accepted to ICLR 2026

点击查看摘要

Abstract:Parameterized and coupled partial differential equations (PDEs) are central to modeling phenomena in science and engineering, yet neural operator methods that address both aspects remain limited. We extend Fourier neural operators (FNOs) with minimal architectural modifications along two directions. For parameterized dynamics, we propose a hypernetwork-based modulation that conditions the operator on physical parameters. For coupled systems, we conduct a systematic exploration of architectural choices, examining how operator components can be adapted to balance shared structure with cross-variable interactions while retaining the efficiency of standard FNOs. Evaluations on benchmark PDEs, including the one-dimensional capacitively coupled plasma equations and the Gray-Scott system, show that our methods achieve up to 55-72% lower errors than strong baselines, demonstrating the effectiveness of principled modulation and systematic design exploration.

[LG-71] Generalization bounds and sample complexity for remaining useful life prediction from complete degradation trajectories

链接: https://arxiv.org/abs/2607.23454
作者: Huy Hoang Le,Kim-Anh Nguyen
类目: Machine Learning (cs.LG)
*备注: This manuscript has been accepted for publication in Measurement Science and Technology. The final Version of Record is available at this https URL

点击查看摘要

Abstract:Data-driven remaining useful life (RUL) prediction requires complete degradation trajectories for training, yet such run-to-failure data are scarce and expensive. Practitioners currently lack principled guidance on how many failure examples suffice for a given model and accuracy target. This paper develops a sample complexity framework for RUL prediction comprising seven main results organised around three themes. First, we establish fundamental learning rates: a distribution-free generalization bound shows that the uniform deviation of the mean squared error decreases as O(B^2\sqrtp/n) , where p is the model complexity and n the number of trajectories, and a minimax lower bound proves that the \Theta(p/n) rate is unimprovable. \revSecond, we quantify how domain knowledge accelerates learning: incorporating degradation physics reduces data requirements by up to two orders of magnitude for deep networks, a Bernstein-type analysis achieves the minimax-optimal O(p/n) rate under high signal-to-noise conditions, and closed-form penalties reveal when an incorrectly assumed physics model hurts rather than helps. Third, we characterise the impact of data quality: fleet variability induces an irreducible bias - variance tradeoff, while right-censored observations suffer an efficiency loss that depends critically on the degradation class. Closed-form expressions are provided for exponential, power-law, and stretched-exponential degradation. \revCross-domain validation against published turbofan, battery, and bearing benchmarks confirms the theoretical predictions within a factor of 2 - 3 on average. The results yield practical guidelines for planning data collection, selecting model complexity, and evaluating physics model assumptions in prognostics applications.

[LG-72] Local Regularization Does Not Characterize Multiclass PAC Learnability

链接: https://arxiv.org/abs/2607.23449
作者: Eric Hou
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Local regularization assigns each hypothesis a test-point-dependent score and predicts with a minimum-score hypothesis consistent with the sample. Asilis et al. asked whether this principle characterizes multiclass PAC learnability. We give a negative answer. There is a countable class of Daniely–Shalev-Shwartz dimension at most two with realizable PAC sample complexity [ O!\left(\frac1\varepsilon\log\frac1\delta\right), ] that no local regularizer learns. Hypotheses are edges of complete graphs and instances are tournaments. At a test tournament, the scores fix an edge ranking while the training sample independently removes competitors. Cyclic triangles force enough inversions that surviving competitors produce constant population error at arbitrarily large sample sizes.

[LG-73] PerturbPFN: Probing the Limits of Synthetic Priors in Drug Perturbation Modelling ICML

链接: https://arxiv.org/abs/2607.23447
作者: Yuche Gao,José Miguel Hernández-Lobato,Siyuan Guo
类目: Machine Learning (cs.LG)
*备注: 19 pages. Accepted at the 2nd ICML Workshop on Foundation Models for Structured Data (FMSD 2026), Seoul, South Korea

点击查看摘要

Abstract:Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional expression responses, and limited experimental coverage of the large small-molecule design space. We propose PerturbPFN, a PFN-style amortized model for unknown-target perturbation prediction under a hierarchical synthetic structural prior. Instead of directly regressing high-dimensional expression responses, PerturbPFN infers a latent system graph, sparse atomic intervention targets, and intervention strengths, then propagates their effects through an SCM decoder. The model is trained entirely on prior-predictive synthetic episodes generated from biologically motivated graph and expression simulators, enabling structured in-context learning without test-time gradient updates. We evaluate PerturbPFN on both real single-cell perturbation data and synthetic benchmarks, covering effect prediction, target identification, and regulatory structure discovery. Our results show that PerturbPFN offers a complementary trade-off to specialized baselines, achieving competitive perturbation prediction with low inference cost while exposing interpretable intermediate estimates of targets, strengths, and system structure.

[LG-74] Neural Representation of Minimal Surfaces

链接: https://arxiv.org/abs/2607.23437
作者: Jiayin Sun,Albert Chern
类目: Graphics (cs.GR); Machine Learning (cs.LG)
*备注: 11 pages, 11 figures

点击查看摘要

Abstract:We propose a neural representation for minimal surfaces. Unlike prior approaches based on discretization or Physics-Informed Neural Networks (PINNs), where meshes or neural fields are optimized to approximate the governing equations, our method builds on an exact representation, similar to the classical Weierstrass–Enneper parameterization, yielding minimal surfaces up to negligible quadrature error in evaluation. We formulate a training objective for the Plateau problem that optimizes over this representation.

[LG-75] Short-Term Pain for Long-Term Gain: Adaptive Experiment with Post-Commitment Reward Shift

链接: https://arxiv.org/abs/2607.23432
作者: Puping Jiang,Wei Tang
类目: Machine Learning (cs.LG)
*备注: A one-page abstract appeared in ACM EC’26

点击查看摘要

Abstract:Decision-makers in learning environments face a dilemma when their short-term optimal actions may not favor their long-term benefits the most. To understand the fundamental tradeoff behind the dilemma, we study adaptive experimentation with post-commitment reward shifts. During an experiment phase, the decision-maker may adaptively test multiple options; during a subsequent commitment phase, the decision-maker must commit to a single option, whose reward may differ from its pre-commitment reward. We propose the Reserved Arm Eliminations for Commitment (RAEC) algorithm, which reserves a predetermined portion of the experiment phase to identify the best post-shift option while using the remaining rounds to minimize short-run regret. We establish regret upper bounds for RAEC across all parameter regimes and matching minimax lower bounds, providing a tight characterization of the cost of balancing short-term performance and long-term commitment. We also study two extensions. With prior structural knowledge linking pre- and post-shift rewards, we show that correctly identifying the ranking-changing component of the shift is more important than estimating its absolute magnitude. For settings with concave commitment rewards and portfolio choice, we develop the Reserved Online Stochastic Convex Optimization for Commitment (ROSCOC) algorithm, which directly converts its reserved exploration history into a commitment portfolio and achieves tight regret bound. Finally, we also conduct numerical experiments which confirm that our proposed algorithms achieve the desired regret predicted by our theory, and also outperform other baseline algorithms. Comments: A one-page abstract appeared in ACM EC’26 Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.23432 [cs.LG] (or arXiv:2607.23432v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.23432 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-76] Harmonized Interpretable ECG Waveform Features for Robust Cross-Dataset Clinical Prediction

链接: https://arxiv.org/abs/2607.23412
作者: Jie Lin,Weijie Sun,Sunil V. Kalmady,Anita Khalafbeigi,Abram Hindle,Padma Kaul,Russell Greiner
类目: Machine Learning (cs.LG)
*备注: Full version of the work presented as a 2-page paper at the 39th IEEE International Symposium on Computer-Based Medical Systems (CBMS 2026)

点击查看摘要

Abstract:Electrocardiograms (ECGs) are widely used for cardiovascular risk prediction, yet models often fail to transfer across hospitals because of protocol, population, and measurement differences. We benchmark cross-dataset generalization on three tasks - heart failure classification, 30-day all-cause mortality, and 30-day mortality among sinus-rhythm ECGs - using two large cohorts (MIMIC-IV and the Alberta Cohort). To reduce vendor-specific measurement mismatch, we build a harmonized, interpretable feature representation computed directly from raw waveforms: FeatureDB morphology/heart-rate-variability summaries plus compact time-frequency descriptors (autoregressive and wavelet features). We train XGBoost models on this unified feature space and evaluate with patient-disjoint internal and bidirectional external testing. We pre-specify two hypotheses: (H1) external AUROC retains at least 90% of source-site internal AUROC under transfer, and (H2) internal AUROC of the harmonized feature set stays within 10% of dataset-native machine-measurement models. Across tasks, internal AUROC is 0.79-0.82 and cross-dataset AUROC is 0.74-0.78, with larger and direction-dependent AUPRC shifts under transfer. As an exploratory benchmark, an end-to-end ConvNeXt model trained directly on raw ECG waveforms with age and sex achieves higher internal AUROC, while the harmonized representation remains competitive in relative cross-dataset transfer stability. These findings show that a consistent waveform-derived feature interface preserves performance, supports realistic external validation, and provides a transparent alternative for cross-site clinical prediction.

[LG-77] ransfer Learning Architectures for Scalable Multi-Fidelity Bayesian Optimization

链接: https://arxiv.org/abs/2607.23404
作者: Jaewook Lee,Ethan Errington,Christian D. Lorenz,Miao Guo
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, expensive ones, with a predictive surrogate at its core. Gaussian processes (GPs) are the default choice, but they scale poorly as data accumulate and assume a smooth landscape that molecular and materials search spaces routinely violate. Transfer learning offers an alternative suited to this regime: it learns a representation from abundant cheap data and adapts it to sparse expensive data. Despite its use in property prediction, transfer learning has not been tested as the engine of a closed-loop optimization. Here we benchmark eleven transfer-learning surrogates against four GP methods under an identical selection rule, fidelity budget, and model size, across nine tasks spanning synthetic functions to real chemistry and materials problems. GPs win on smooth, low-dimensional functions but perform worst on molecular and materials problems, where transfer-learning surrogates reach substantially better solutions using far less computation. Because acquisition policy is held fixed across surrogates, this advantage is attributable to the surrogate itself. Uncertainty-driven exploration is not reliably beneficial, and calibration does not predict optimization performance, so greedy exploitation of the transfer-learned mean is the more robust default. Transfer learning is therefore the surrogate of choice for molecular and materials MFBO.

[LG-78] A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks

链接: https://arxiv.org/abs/2607.23397
作者: Sumio Watanabe
类目: Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood. In the infinite-width limit, two different theoretical frameworks have been proposed. One reduces deep learning to kernel regression with a fixed kernel by assuming that the parameters remain close to their initialization, whereas the other allows the parameters to move away from their initialization, requiring the kernel itself to be optimized. In this paper, we study a three-layer neural network with a finite but large number of hidden units. We show that training the input-to-hidden weights yields a smaller generalization error than keeping them fixed. Furthermore, the latter setting exhibits singularities in the parameter space, whereas the former does not. These findings indicate that singularities play an essential role even in wide neural networks. Subjects: Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML) Cite as: arXiv:2607.23397 [cs.LG] (or arXiv:2607.23397v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.23397 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-79] Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

链接: https://arxiv.org/abs/2607.23395
作者: Roman Solovyev,Ilya Kiselev,Alexander Stempkovskiy,Tatiana Gabruseva
类目: ound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
*备注:

点击查看摘要

Abstract:Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production. The separation quality depends on engineering decisions across the entire pipeline: model choice, training data preparation and augmentation, loss function and metrics choice, training configuration, validation, and post-processing. This paper presents MSST (Music-Source-Separation-Training) - a universal open-source framework for MSS tasks, which unifies training, validation, and inference for a broad range of modern demixing model families under a single, configuration-driven interface. The framework supports various model architectures, data preprocessing and augmentations, multiple loss functions and evaluation metrics, which helps with fast iterations and ablation studies. Additionally, the framework supports a range of practical techniques that improve separation quality, such as sliding-window inference with cross-fading, test-time augmentation, model ensembling, and fine-tuning via Low-Rank Adaptation (LORA). Our ablation studies demonstrate improvements of MSS using the above techniques. By consolidating these components into a reproducible, YAML-configurable framework, MSST lowers the barrier to systematic experimentation and enables rapid iteration from idea to verifiable result.

[LG-80] When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation

链接: https://arxiv.org/abs/2607.23390
作者: Mojtaba Soltanalian
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注: 141 pages, 26 figures, 15 tables. Includes complete proofs and documents the QReplace decision-support and Lean 4 verification companions. To be submitted to the Journal of Machine Learning Research

点击查看摘要

Abstract:When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map? We model a quantized residual system over a fixed horizon as a pure schedule selecting fields from a declared low-bit operation library, and use relaxed controls to characterize its infinite-depth limit. The distance from the target to the closed relaxed reachable set is the exact structural floor: no increase in depth can remove it for that library. Pure schedules approach the relaxed class at rate O(D^-1) under bounded-variation time dependence and O(D^-\vartheta+D^-1) under Holder dependence of exponent \vartheta . Execution arithmetic can reverse this conclusion: full-state write-back introduces a D\rho_z penalty and can freeze residual updates, whereas increment error feedback replaces this growth by a bounded carry term and obeys an exact common-lattice conservation law. A fixed-teacher converse makes this rate sharp: for coherent depth- L first-order high-precision comparators, accuracy matching requires D=\Theta(L) . Learned codebooks add a metadata resource, while state-dependent routing introduces hybrid event conditions. Verified primal and dual bounds yield feasible, impossible, or unresolved decisions before training. Companion software implements the workflow, and Lean 4 machine-checks the exact discrete core. Depth replaces precision only relative to a declared library, horizon, execution semantics, and routing model.

[LG-81] Rendering on Real Silicon: GPU Render-Timing as a Passive AI-Resistant CAPTCHA Signal

链接: https://arxiv.org/abs/2607.23389
作者: David Noever,Forrest McKee
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Conventional CAPTCHAs pose puzzles that modern AI systems increasingly solve, while behavioral and cryptographic-attestation defenses carry privacy or enrollment costs. We investigate an orthogonal signal: the physical timing behavior of a client’s GPU under a controlled WebGL rendering workload. Unlike WebGL fingerprinting, which hashes pixel output into a static device identifier, we measure render-timing dynamics to classify rather than identify, leaking no persistent identifier. We characterize the in-the-wild adversary with a 12-hour passive deployment (207 unsolicited requests; 86% automated; 85% of browser-claiming clients failed HTTP header-consistency checks). We then collect labeled GPU-timing samples through a single public endpoint exercised by real browsers (positive class, 13 distinct GPUs) and by keyed headless automation across a render-backend matrix (negative class). Software-rendered automation – empirically the dominant real-world adversary – separates from genuine GPUs by roughly 5x in mean render time. On a confound-controlled comparison (identical GPU family and browser engine, differing only in headless vs. interactive execution), headless automation on real hardware still exhibits a distinct timing signature, separating from human samples by 75-106% on frame jitter, timer-quantization ratio, and coefficient of variation. We report these as pilot-scale findings on a single GPU architecture and outline the cross-architecture collection required to establish generalization.

[LG-82] Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features

链接: https://arxiv.org/abs/2607.23370
作者: Muhammad Abdullah Haroon
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE); Econometrics (econ.EM); Computation (stat.CO)
*备注: 19 pages, 15 figures, 4 tables

点击查看摘要

Abstract:Bitcoin price prediction on sub-daily timescales is a hard open problem in computational finance. Bitcoin exhibits fat-tailed returns, non-stationary dynamics, and a price discovery process influenced by social discourse on Reddit and Twitter. Conventional approaches fuse OHLCV technical features with sentiment via static concatenation, applying identical fusion weights regardless of market state. This is inconsistent with the behavioural finance literature, which shows that retail sentiment is most predictive during volatile periods and noisy during calm ones. This paper proposes Regime-Aware Multi-Modal Learning (RAML), which conditions fusion of sentiment and price features on a dynamically detected binary market regime. Rolling 24-hour volatility partitions observations into stable and volatile regimes; a learnable sigmoid gate adjusts the weight of the sentiment embedding relative to the price embedding, trusting sentiment more during volatility and price dynamics more during stable phases. The system is evaluated on 3,491 hourly observations (July 2024-September 2025), combining Bitcoin OHLCV data with Reddit /r/Bitcoin FinBERT sentiment. Four models are compared - price-only BiLSTM, sentiment-only classifier, static-concatenation BiLSTM, and RAML - across 3-hour and 6-hour horizons, with an ablation study isolating the sentiment branch, regime detection, and adaptive fusion. RAML achieves macro-F1 of 0.5474 (3h) and 0.5513 (6h), with the highest AUC at 3 hours (0.5084), indicating better calibration. Ablation confirms every component is necessary, and replacing adaptive weighting with concatenation causes recall collapse at 6 hours (F1: 0.14). These results establish regime-conditioned adaptive fusion as a necessary design principle for multi-modal financial forecasting.

[LG-83] On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards

链接: https://arxiv.org/abs/2607.23364
作者: Fei Ding,Yongkang Zhang,Yuhao Liao,Zijian Zeng,Huiming Yang
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1. The recent improvement Dr. GRPO (COLM 2025) identifies the response-level length bias caused by per-trajectory length normalization in GRPO and proposes removing this normalization, claiming the resulting optimizer is “unbiased.” We show that this claim is incomplete. Specifically, we establish an impossibility theorem: under the standard outcome reward + GRPO setting, no length-based weighting scheme can simultaneously achieve the following two properties. (P1) Gradient unbiasedness: the gradient estimator is an unbiased estimate of the true policy gradient. (P2) Length invariance: each trajectory’s effective contribution to the gradient is independent of its token length. GRPO approximately satisfies P2 but violates P1; Dr. GRPO satisfies P1 but violates P2. We characterize the complete tradeoff spectrum via the parametric family f_alpha(L) = L^alpha - 1, where alpha = 0 recovers GRPO, alpha = 1 recovers Dr. GRPO, and provide quantitative analysis showing that Dr. GRPO’s length bias can cause longer trajectories to dominate gradient updates by a factor proportional to the length ratio. Our results reveal that neither algorithm is universally “done right”; they occupy opposite ends of a fundamental and unavoidable tradeoff.

[LG-84] Exploration of the generative capabilities of Boltzmann machines applied to social systems under the majority rule

链接: https://arxiv.org/abs/2607.23349
作者: Mauricio A. Valle,Gonzalo A. Ruz
类目: Machine Learning (cs.LG); Computers and Society (cs.CY)
*备注:

点击查看摘要

Abstract:We study the generative capabilities of Boltzmann machines to recover systems governed by the majority rule under critical conditions. To this end, we train deep belief networks (DBNs) with different configurations, where the first layer can use Gaussian visible units with more than two states (i.e., non-binary units). We then allow the DBN to “dream” samples conditioned on visible units that we keep fixed, and we measure the deviation of this dreamed system from the real one. We also corroborate, using a discrete thermometer based on a convolutional network, that the reconstructions remain in a critical state. Across several training sessions with different architectures, we show that, despite the complexity of the problem, the DBN can recover samples that remain critical even under input noise, with a gradual degradation of physical observables relative to the original sample.

[LG-85] SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation

链接: https://arxiv.org/abs/2607.23346
作者: Aditya Dewan,Arjun Yogeswaran,Benjamin Fedoruk
类目: Machine Learning (cs.LG)
*备注: 12 pages, 5 figures. Code and reproduction scripts: this https URL . Awarded 2nd Place in Mathematical and Cybersecurity Research (NSA) at Regeneron ISEF 2023 and an Outstanding Research Award at WAICY (World AI Competition for Youth) 2022

点击查看摘要

Abstract:Modern deep neural networks are potent catalysts for scientific and industrial impact, yet excessive parameter counts impede deployment in low-compute settings such as hospital equipment and energy infrastructure. Predominant knowledge distillation (KD) methods favor replication: smaller students mimic teacher output logits, yet empirically yield low task performance, hamper convergence, and act merely as regularization rather than substantive knowledge transfer. We propose Saddle Point Recruitment for Knowledge Distillation (SPRKD), reframing distillation from replication to employing teachers as optimization-curvature and domain proxies, characterizing saddle points as regions of strong further-descent potential via embedding and basin-fractal properties. Using Hessian eigenvalue spectral density (ESD), SPRKD identifies low-loss saddle regions for student re-exploration; weak-teacher ensembles are aggregated into an Approximated Saddle Region (ASR), re-parameterized into the student via Transfer Learning by Injection, and approached with exponentially decaying Euclidean transformations, Negative Hessian Eigensteps, and Gaussian perturbations. On malaria blood smear classification with a 6,430-parameter CNN distilled from a weak 25,546-parameter teacher, SPRKD reaches 94.8% validation accuracy, outperforming Response KD by 24.70 percentage points (McNemar p = 6.3e-87) and matching scratch-trained baselines of the same architecture to statistical equivalence (p = 1.0). Across MNIST, CIFAR-100, and TinyImageNet, SPRKD exceeds scratch-trained baselines by up to 8 percentage points on preliminary benchmarks. Hessian ESD and 2-D loss landscape analysis show convergence to wider minima with substantially smaller Hessian trace and spectral radius than Response KD and control students, indicating smoother descent and greater noise robustness.

[LG-86] Does Graph Compression Preserve Signal Propagation?

链接: https://arxiv.org/abs/2607.23338
作者: Kawshik Banerjee,Khaled Mohammed Saifuddin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Graph compression reduces the computational cost of graph learning, but its effect on signal propagation remains largely underexplored. Existing work evaluates compression through downstream task performance or structural preservation, neither of which directly captures how propagation dynamics change after compression. We study two fundamental compression paradigms, coarsening and sparsification, and ask whether they preserve the propagation behavior of the original graph. Across five datasets, varying compression rates, and propagation depths, we measure signal behavior through three complementary metrics. Our results reveal a consistent tension between the two compression families. Sparsification retains higher signal diversity and mitigates oversmoothing, but its propagation trajectory progressively diverges from that of the original graph. Coarsening more faithfully preserves propagation behavior, but at the cost of stronger smoothing and rank collapse. These findings demonstrate that two propagation-centric objectives, preserving signal diversity and preserving propagation fidelity, are distinct and empirically at odds under graph compression, highlighting the need for evaluation protocols that jointly consider both dimensions. The code and results are available at: this https URL

[LG-87] Neural operator discovery from heterogeneous trajectories

链接: https://arxiv.org/abs/2607.23337
作者: Zituo Chen,Qiaofeng Li,Jiaxin Hu,Sili Deng
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Neural operators provide data-driven mappings for modeling dynamical systems. Extending them to families of systems typically requires explicit conditioning variables such as physical parameters, geometries, or boundary conditions. In many real-world settings, these quantities are unobserved. Here, we formulate neural operator discovery (NOD) as the problem of learning both shared solution operators and system-specific variation directly from heterogeneous trajectories without access to labeled governing factors. We introduce a factorized latent-conditioning formulation that jointly learns a neural operator and a low-dimensional latent representation through factorized prediction, trajectory-decoupled sampling, and dimension selection. Across diverse systems, the learned latent representation captures the intrinsic dimensionality of system variation and organizes system instances in a smooth and approximately invertible latent structure aligned with the underlying governing factors. This organization enables generalization to previously unseen system instances, including zero-shot extrapolation across regimes and stable long-horizon prediction. These results establish an interpretable paradigm for operator learning in the absence of explicit factor supervision.

[LG-88] AlloBench: Measuring Online Tool Allocation Capability in LLM Agents

链接: https://arxiv.org/abs/2607.23332
作者: Daniel Wang,Andrew Xu
类目: Machine Learning (cs.LG)
*备注: 24 pages, 6 figures, 8 tables

点击查看摘要

Abstract:Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer an agent that creates a small number of highly reusable tools, rather than many one-offs. We introduce a paired benchmark that tests whether LLM agents exhibit conscious allocation behavior under a fixed budget in two contexts: an abstract text-based formulation and a code-construction task. We find that every frontier model we test—Claude Haiku, Claude Opus, GPT-5.4-mini, and GPT-5.6 Sol—acts near-optimally in the abstract framing but fails to transfer this ability to script-writing. Through further experiments, we identify the particular failure modes for each model. Notably, the first three models fail even when the scripts are not evaluated, while GPT-5.6 Sol stays selective under that weaker manipulation and collapses only at full construction. Furthermore, an open-source Qwen model policy-trained for abstract allocation generalizes this ability across held-out lexical variations, but sees no improvement at script allocation. Together, these results establish online tool allocation as a significant capability boundary, even for modern frontier models.

[LG-89] StageGuard: Physiologically Constrained Sleep Staging KDD KDD2026

链接: https://arxiv.org/abs/2607.23284
作者: Juntang Wang,Yihan Wang,Hao Wu,Jiayu Gao,Shixin Xu,Dongmian Zou
类目: Machine Learning (cs.LG)
*备注: 12 pages. Accepted at KDD 2026 (32nd ACM SIGKDD Conference), AI for Sciences track

点击查看摘要

Abstract:Automated sleep staging is increasingly used in large-scale studies to derive sleep-architecture endpoints: total sleep time, REM latency, sleep efficiency, and bout-duration statistics. Deep learning models achieve epoch-level accuracy approaching inter-rater agreement, yet often produce hypnograms that violate physiological invariants, such as rare transitions (e.g., direct Wake - REM) or excessively fragmented sequences. Such violations can bias downstream sleep metrics, regardless of overall accuracy. We propose StageGuard, a plug-and-play, backbone-agnostic structured-inference framework that wraps any neural sleep-staging backbone with physiology-informed priors. StageGuard combines (1) a differentiable soft transition penalty that discourages physiologically rare transitions during training, and (2) a semi-Markov constrained decoder with a duration-augmented state space that jointly enforces transition penalties and minimum bout durations at inference. Unlike hard-prohibition methods, it admits rare transitions when emission evidence is overwhelming, leaving informative pathological events recoverable rather than blocked. StageGuard constrains staging outputs to satisfy known physiological priors rather than modeling sleep generatively. We quantify the validity gap using transition-violation rate (TVR) and fragmentation index (FI) and demonstrate that, across six backbones and four datasets, StageGuard reduces TVR to physiologically plausible levels and lowers FI by 56-62%, while maintaining or slightly improving classification accuracy. Crucially, improved constraint satisfaction translates into 59-79% lower error on derived sleep-architecture statistics not directly optimized by the method, and recovers the direction and effect size of expert-defined subgroup differences (OSA severity, age) more faithfully than the unconstrained baseline.

[LG-90] FedTaste: Topology-Aware Structural Transfer for Multimodal Federated Learning with Missing Modalities ACM-MM

链接: https://arxiv.org/abs/2607.23245
作者: Haochen Liang,Jie Zhang,Hideya Ochiai
类目: Multimedia (cs.MM); Machine Learning (cs.LG)
*备注: Accepted to ACM Multimedia (ACM MM) 2026

点击查看摘要

Abstract:Multimodal Federated Learning is often challenged by arbitrary modality missingness and Non-IID data distributions, which lead to severe representation drift and hinder effective collaboration across clients. Existing methods typically rely on generative imputation, external auxiliary data, or isolated unimodal training to bridge modality gaps, often incurring substantial communication and computational costs as well as potential privacy risks. To address these limitations, we propose FedTaste, a parameter-efficient framework for topology-aware structural transfer in Multimodal Federated Learning with missing modalities. Instead of aligning fragile first-order features, FedTaste focuses on more stable group-level semantic relations. Specifically, FedTaste leverages frozen foundation models to extract a joint multimodal topology from full-modality clients, which is then consolidated by the server into a global structural blueprint. To adapt clients with missing modalities, we introduce Modality-Adaptive Structural Prompts together with spectral consistency regularization, enabling lightweight branch-specific adaptation that aligns local partial representations with the shared blueprint. In this way, FedTaste avoids explicit modality imputation while preserving shared semantic structure across clients. Extensive experiments demonstrate that FedTaste consistently achieves superior performance across multiple datasets and challenging Non-IID settings, while substantially reducing communication overhead compared with existing methods.

[LG-91] From Score Learning to Discretized Sampling: An End-to-End Generalization Analysis of Diffusion Models

链接: https://arxiv.org/abs/2607.23226
作者: Jinshu Huang,Yiming Jiang,Chunlin Wu
类目: Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:

点击查看摘要

Abstract:Despite the empirical success of score-based diffusion models, a complete theoretical understanding of how finite-sample learning, network parameterization, and numerical discretization jointly dictate generative quality remains underdeveloped. Existing sampling analyses often evaluate the generative performance conditional on an oracle score or a pre-specified error threshold. In this work, we establish a unified convergence and generalization framework for score-based diffusion models parameterized by practical ResNet-type architectures. We analyze the generalization and convergence properties from the practical finite-sample, discrete-time learning problem of the score function to the ideal continuous-time, population-level objective. Based on the generalization result of the learning problem of score function, we analyze the sampling process induced by the learned score function and provide an end-to-end total variation distance estimate for the generated terminal distribution. This estimate explicitly decomposes the overall generative error into four interpretable components: the truncation error of the forward process, the reverse-time discretization error, the generalization error incorporating both finite data and forward-time discretization, and the training optimization gap. Our results quantitatively characterize how the training sample size, temporal discretization grids, and optimization accuracy jointly control the final fidelity of samples generated by diffusion models.

[LG-92] ParasGB: A Graph Benchmark Suite for Parasitic Estimation on AMS Circuits

链接: https://arxiv.org/abs/2607.23225
作者: Jiajun Zou,Jiawei Liu,Ao Liu,Junnong Tian,Yibin Zhang,Chengjie Liu,Yuxi Wang,Shan Shen,Wenhua Gu,Jun Yang,Wenjian Yu
类目: Machine Learning (cs.LG)
*备注: Published at ICCAD2026. Full appendix version

点击查看摘要

Abstract:As chip manufacturing processes advance to deep submicron nodes, parasitic interconnect effects increasingly dominate the performance of analog and mixed-signal (AMS) circuits and often lead to costly layout iterations. This makes early-stage estimation of parasitic capacitance and resistance important for parasitic-aware design exploration before full physical implementation. However, progress on GNN-based parasitic modeling has been hindered by the lack of public, high-fidelity RC benchmarks that support reproducible evaluation. To address this gap, we introduce ParasGB, the first open-source benchmark suite for pre-layout parasitic parameter prediction on circuit graphs. ParasGB provides large-scale, heterogeneous RC networks extracted with commercial EDA tools from tape-out-proven designs, together with a unified evaluation protocol covering node-level ground capacitance, edge-level resistance, and edge-level coupling capacitance. Within this framework, we benchmark diverse GNN architectures using a standardized training pipeline and expose challenges such as extreme label imbalance, long-tailed parasitic distributions, and strong structural heterogeneity. By establishing a physically grounded and standardized benchmark for early-stage parasitic prediction, ParasGB provides an open platform for reproducible research on circuit graph learning and parasitic-aware model development. All datasets, preprocessing scripts, and configurations are publicly available in our code repository this https URL.

[LG-93] Variance-Preserving Orthogonal Selection (VPOS): Greedy Feature Selection via Orthogonal Deflation in PCA Loading Space

链接: https://arxiv.org/abs/2607.23198
作者: Baran Koseoglu,Berrin Yanikoglu
类目: Machine Learning (cs.LG)
*备注: 23 pages

点击查看摘要

Abstract:We propose Variance-Preserving Orthogonal Selection (VPOS), a greedy framework for unsupervised feature selection that operates in the weighted PCA loading space. After each selection, VPOS projects out the chosen feature’s variance direction via null-space deflation, forcing subsequent selections to cover orthogonal parts of the covariance structure. Each step provably reduces the loading matrix rank by one, and the greedy objective connects to monotone submodular maximization. The single hyperparameter d is selected via a reproducible rule: the value minimising reconstruction MSE in a sensitivity sweep. On eight benchmarks, VPOS achieves the lowest reconstruction MSE on all eight while running 10-140x faster than graph-based methods at scale. Comparing against PCA (no deflation) at matched d confirms deflation as the primary driver, reducing MSE by 10-73%.

[LG-94] Domain-Prior-Regularized Graph Modeling for Anomaly Detection in Cyber-Physical Systems ICML2026

链接: https://arxiv.org/abs/2607.23197
作者: Youngseok Hwang,Joonsung Kwon,Geonwoo Lee,Hyunwoo Park
类目: Machine Learning (cs.LG)
*备注: 12 pages, ICML 2026 AI for Science Workshop

点击查看摘要

Abstract:Anomaly detection on multivariate sensor time series is critical for industrial monitoring of cyber-physical systems (CPS), where even subtle deviations from normal behavior can indicate process disruption. Recent graph-based approaches have made significant progress, but they often struggle in small-scale physical systems with scarce labeled anomalies and limited normal data. In such settings, graph-based models tend to capture spurious correlations and produce unstable sensor topologies. We propose DPR-GM (Domain-Prior-Regularized Graph Modeling), a forecasting-based framework that incorporates system design knowledge into graph construction. DPR-GM leverages a large language model (LLM) to extract directed physical couplings between sensor pairs from system documentation, which are encoded as a binary domain adjacency matrix serving as a structural gate over sensor relations. This gate is then modulated by Pearson correlations estimated from normal training data. The anomaly score is further weighted by sensor-level reliability derived from the coefficient of variation. All graph and weighting components are fixed prior to training and add no learnable parameters. On the SKAB benchmark, DPR-GM outperforms graph-based, statistical, and deep learning baselines across F1, AUROC, and AUPRC, showing that domain-structured graph priors are a practical alternative to fully learned topologies in data-scarce CPS.

[LG-95] Data-Driven Diffusion Processes on Differential Forms via the Projected Ambient Connection Laplacian

链接: https://arxiv.org/abs/2607.23192
作者: Alvaro Almeida Gomez,Jorge Duque Franco
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG)
*备注: Comments are welcome

点击查看摘要

Abstract:We develop a data-driven approximation of the projected ambient connection Laplacian acting on differential forms over smooth Riemannian manifolds sampled by point clouds. The proposed construction extends the classical framework of diffusion maps and Vector Diffusion Maps from scalar functions and tangent vector fields to differential forms of arbitrary degree. Our approach is based on a novel representation of differential forms as alternating differential arrays obtained through an extension of the classical musical isomorphism. This representation enables the construction of a matrix-valued diffusion operator that approximates the projected ambient connection Laplacian directly from point cloud data without requiring a mesh or simplicial complex. The proposed discretization admits the asymptotically optimal kernel bandwidth scaling inherited from diffusion maps, leading to sharper convergence guarantees than previous data-driven approximations of the Hodge Laplacian. Building upon this operator, we derive a fully data-driven explicit Euler scheme for the heat equation on differential forms and validate the proposed methodology through numerical experiments on the unit sphere. The experiments confirm the predicted decay of the analytical solution and demonstrate the effectiveness of the proposed discretization. The proposed framework provides a natural generalization of Vector Diffusion Maps to differential forms of arbitrary degree and establishes a practical foundation for the numerical approximation of geometric partial differential equations directly from point cloud data. Comments: Comments are welcome Subjects: Numerical Analysis (math.NA); Machine Learning (cs.LG) Cite as: arXiv:2607.23192 [math.NA] (or arXiv:2607.23192v1 [math.NA] for this version) https://doi.org/10.48550/arXiv.2607.23192 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Alvaro Almeida Gomez [view email] [v1] Sat, 25 Jul 2026 13:14:46 UTC (354 KB)

[LG-96] Wrong Design Intent Is Worse Than None: A Derangement-Control Diagnosis of Header Conditioning in CAD Program Completion

链接: https://arxiv.org/abs/2607.23191
作者: Yang Xiao
类目: Machine Learning (cs.LG)
*备注: 19 pages, 4 figures

点击查看摘要

Abstract:Fine-tuned code LLMs can be conditioned on a lightweight design-intent header to steer parametric CAD generation, but whether the model actually reads the header’s content has not been tested under a metric independent of the conditioning itself, nor with a causal control. We study CADCON, a five-feature design-intent header prepended to CadQuery-style sketch-extrude programs during LoRA fine-tuning of Qwen2.5-Coder-1.5B, re-scored by executable geometric assertions on the produced B-rep solid, sharing no code with the header-defining regex extractor. Across three seeds and a pre-registered 0%, 40%-prefix \times correct, wrong, masked-header matrix we find: (i) in conditional completion (40% prefix), a semantically wrong header degrades adherence below the no-header baseline (0.43 \to 0.30/0.21 text/token) on design intents the model can render unconditioned – polygonal and thin geometries; circle and tall intents sit at a baseline generation floor for this checkpoint ( \approx 0 in both compared arms) and are uninformative for this contrast; (ii) a derangement control – retrained with shuffled ground-truth headers, identical header marginal but destroyed content correlation – remains competent yet is immune to wrong headers, while the standard model is not (text headers; interaction significant on 3/3 seeds, p \leq 4.2\times10^-3 ): the harm requires the learned header \to program mapping, excluding the marginal/mechanical distribution-shift confound; (iii) the independent metric deflates the apparent benefit of a correct header (token: +0.21 regex \to +0.02 geometry), quantifying metric circularity; (iv) the harm is regime-specific – at 0% prefix the unconditioned baseline cannot generate valid CAD at all. Wrong intent is not noise: it actively misdirects generation.

[LG-97] XGRVFL-MV: Residual-Coupled Graph-Embedded Multi-View Random Vector Functional Link Network with FleXi Guardian Loss

链接: https://arxiv.org/abs/2607.23149
作者: Yogesh Kumar,Mudasir Ganaie
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Random Vector Functional Link (RVFL) networks provide an efficient randomized learning framework for classification. Existing multi-view RVFL methods utilize complementary information from multiple views. However, preserving view-specific geometric structure, limiting the influence of large prediction residuals, and modeling relationships between multiple views remain challenging. This paper proposes a Residual-Coupled Graph-Embedded Multi-View RVFL model with fleXi guardian loss (XGRVFL-MV) for multi-view classification. The proposed model constructs RVFL representation for each view, incorporates graph embedding with intrinsic and penalty graphs constructed using the Local Fisher Discriminant Analysis weighting scheme. It also uses the bounded and asymmetric FleXi Guardian (XG) loss for residual learning. A residual-coupling term is introduced to encourage consistency among view-specific prediction residuals while preserving view-specific representations. The resulting optimization problem is solved using an inversion-free first-order optimization procedure based on Nesterov accelerated gradient descent. We evaluate the proposed model on UCI, KEEL, AwA, and Corel5k benchmark datasets. Experimental results, together with statistical analyses and hyperparameter sensitivity analyses, show that XGRVFL-MV achieves competitive classification performance compared with the baseline methods across the evaluated benchmark datasets.

[LG-98] Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting

链接: https://arxiv.org/abs/2607.23146
作者: Morad Laglil,Bertrand Pracca,Emilie Devijver,Eric Gaussier
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Inspired by recent breakthroughs in large language models for natural language processing, foundation models have emerged as a promising paradigm for zero-shot time series forecasting, enabling accurate predictions on datasets never seen during pre-training. Ranging from tens to hundreds of millions of parameters, these models are pre-trained on vast and diverse collections of time series, learning generalizable representations that support both point and probabilistic forecasting. This approach alleviates the need for dataset-specific model design and manual tuning, offering a unified solution across forecasting problems. In this work, we review the main architectures, pre-training strategies, and optimization methods underpinning these models. We further investigate post-pre-training fine-tuning of selected foundation models to enhance their performance on specific datasets. Our empirical results demonstrate that this step consistently improves forecasting accuracy over the zero-shot baseline.

[LG-99] Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems

链接: https://arxiv.org/abs/2607.23134
作者: Tanmay Khandait,Preetom Biswas,Hideki Okamoto,Bardh Hoxha,Georgios Fainekos,Giulia Pedrielli
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注:

点击查看摘要

Abstract:Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing falsification approaches rely on conditional sampling strategies that factor the joint distribution over environments and system executions, and therefore suffer from multiplicative rarity effects: the simultaneous scarcity of failure-inducing inputs and failure-inducing traces makes exhaustive search prohibitively expensive. This paper develops DiffTilt, a distributional framework that exponentially tilts a diffusion model-induced joint distribution over environments and executions. We show that diffusion-guided sampling admits an exact interpretation as importance sampling in the joint space, where guidance scores induce a KL-optimal reallocation of probability mass towards failure-relevant behaviors. We further show that tilting provably amplifies failure probability and strictly outperforms conditional sampling, which is limited by multiplicative rarity. In this framework, the joint generative model serves as a reusable prior over scenarios and need not faithfully represent the system under test. Expensive system simulations are instead limited to learning a scoring function that characterizes scenario quality, enabling their selective and adaptive use. We study DiffTilt on ARCH-COMP benchmarks, and we propose an additional tractor-trailer benchmark showing the behavior of several approaches when scenario generation is guided by a well-defined specification rather than a reward. The proposed method achieves competitive or improved falsification performance compared to state-of-the-art approaches, with larger gains when specification definition is not limited to STL formulas. Subjects: Machine Learning (cs.LG); Systems and Control (eess.SY) Cite as: arXiv:2607.23134 [cs.LG] (or arXiv:2607.23134v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.23134 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-100] Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation

链接: https://arxiv.org/abs/2607.23125
作者: Shuai Wang,Daoan Zhang,Zhe Tang,Hao Cheng,Jiaheng Wei
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annotated data, distillation from external models, reinforcement learning with human feedback, or verifiable answers. This limits their ability to improve without external supervision. To tackle this, we propose NOPD (Noisy Student On-Policy Self-Distillation), a simple yet effective self-distillation approach that improves VLMs without any external models or ground-truth answers. Our key insight is that prediction discrepancies between clean and corrupted inputs naturally induce a self-supervision signal. In NOPD, the model learns from corrupted inputs while using its own predictions under clean inputs as token-level supervision. We show the effectiveness of NOPD on five visual reasoning tasks; it can match and even outperform reinforcement learning approaches or distillation from external models. Notably, when trained with 2.1K samples from Geometry3K, NOPD improves Qwen2.5-VL-7B by 20 points on its validation set. It also shows generalization on out-of-distribution test sets and achieves 7.4 point gains on MathVista. Furthermore, we demonstrate that NOPD is a general approach to enhance VLMs, achieving improvements across three models on 12 benchmarks.

[LG-101] Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs

链接: https://arxiv.org/abs/2607.23115
作者: Zhihao Xu,Hao Zhong,Zeting Zhou,Yuhang Xu,Haoyu Tong,Wei Wang,Jinshan Chen,Keqiang He,Chong Zhu,Shengzhong Liu,Fan Wu,Guihai Chen
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注: 20 pages, 28 figures

点击查看摘要

Abstract:This paper aims to enable computation- and communication-efficient GPU sharing across devices within local area networks (LANs), facilitating ubiquitous AI inference on heterogeneous personal devices. We achieve distributed task offloading via CUDA API remoting. However, beyond raw computation, network constraints emerge as the primary bottleneck: limited bandwidth, high-frequency API invocations, and cross-task contention significantly hinder performance. To address these challenges, we propose Gleam, a novel and network-efficient framework for task-generic GPU sharing across local-area CUDA devices, with three key contributions. First, we reduce bandwidth overhead in CUDA API remoting through automatic model weight caching, and mitigate accumulated latency from frequent API calls by asynchronous execution. Second, we design a runtime task scheduler that dynamically determines API remoting pairs between LAN clients and servers, explicitly accounting for both network conditions and GPU resource contention under parallel workloads. Finally, we introduce dedicated mechanisms to ensure CUDA context consistency across distributed executions. Extensive experiments on heterogeneous NVIDIA GPUs and diverse AI workloads show Gleam consistently outperforms state-of-the-art baselines, achieving 1.4-24.2 times improvements in API remoting efficiency and up to 1.79 times higher system throughput.

[LG-102] he Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers

链接: https://arxiv.org/abs/2607.23050
作者: Byeong Hoon Yoon
类目: Machine Learning (cs.LG)
*备注: 11 pages, 1 figure, 5 tables. Code and reproduction scripts included as supplementary material

点击查看摘要

Abstract:Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the minimum model capacity required to solve it? We study this through the Entropic Bound, a spectral notion of task-intrinsic capacity for Transformers. We first prove that, in a linear attention surrogate, the intrinsic rank r^* of the token-mixing operator is a tight lower bound: any rank-deficient model incurs unavoidable excess risk, and the bound is achievable at r^* . We further show that gradient descent recovers this rank under standard low-rank implicit-bias assumptions, confirm all three properties empirically, and show r^* is recoverable from data before training. We then ask whether this transfers to real attention. A naive transfer fails, and a controlled interpolation ladder localizes the cause precisely: it is not softmax and not a rank constraint, but the input-conditioned nature of attention’s mixing operator, which a static weight kernel cannot summarize. Motivated by this, we introduce an attention-native intrinsic rank – the minimum query-key kernel rank realizing the task within the attention class – and show that under this definition the full Entropic Bound structure (deficiency, achievability, recovery) is restored for both linear and softmax attention, with the energy effective rank as the estimator robust to softmax distortion. Finally, we map the boundary of data-only predictability: r^* is exactly recoverable for linear QK attention, even without the value map at scale, while softmax attention admits only partial pre-training recovery due to nonlinear inversion and kernel-value identifiability effects. Our results reframe the Entropic Bound from a post-hoc descriptor into an attention-native capacity measure with a precisely characterized predictability frontier.

[LG-103] Online Policy Evaluation for MDPs with Dynamic UBSR Measures

链接: https://arxiv.org/abs/2607.23030
作者: Weikai Wang,Erick Delage
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. Existing approaches either focus on restrictive classes of risk measures or rely on access to a simulator, limiting their applicability in fully online settings. In this work, we propose computationally efficient online learning algorithms for policy evaluation in Markov decision processes (MDPs) with dynamic utility-based shortfall risk (UBSR) measures under linear function approximation. Specifically, we introduce the UBSR-TD algorithm, establish conditions under which it converges almost surely, and develop several variants designed to accelerate convergence. Our formulation shows that existing policy evaluation algorithms for risk-neutral MDPs can be readily adapted to dynamic UBSR settings by incorporating a loss function into the temporal-difference error. Numerical experiments support our theoretical findings, and an application to a perishable inventory management problem with shelf-life uncertainty demonstrates the practical effectiveness of the proposed methods.

[LG-104] Mini-batch Noise Lowers Sharpness via Dominant-Subspace Fluctuations

链接: https://arxiv.org/abs/2607.23012
作者: Junho So,Dongwook Shin
类目: Machine Learning (cs.LG)
*备注: 20 pages, 7 figures

点击查看摘要

Abstract:During SGD training, the gradients often align strongly with the dominant subspace spanned by the top- k eigenvectors of the Hessian of the loss. While this seems to naturally imply that loss reduction mainly occurs within this space, prior work has shown that updates within this dominant subspace make no meaningful progress in reducing the loss. In this work, we argue that the dominant subspace is better understood not as the main space for loss reduction, but as a key subspace for explaining the sharpness dynamics of mini-batch SGD. To explain the role of the dominant subspace in reducing top- k sharpness, we show how the averaged gradient over fluctuations in the dominant directions produces a sharpness correction term, and derive a sharpness correction term induced by mini-batch noise in the dominant directions. Experimental results show that adding the derived correction term to GD brings the sharpness evolution of GD closer to that of SGD.

[LG-105] Recycling computational processes of dynamic programming for combinatorial optimization problems: a reservoir computing approach

链接: https://arxiv.org/abs/2607.23009
作者: Sora Todaka,Akihiro Yamamoto,Nozomi Akashi
类目: Machine Learning (cs.LG)
*备注: 20 pages

点击查看摘要

Abstract:Reusing previously computed results is a long-standing principle for reducing computational cost, but such reuse has largely been confined to a single problem’s computation. Sharing computational processes across multiple simultaneously solved problems remains possible in principle, yet designing algorithms that exploit nontrivial cross-task relationships is difficult to do manually. Here, we use machine learning to discover such algorithms automatically. Specifically, based on reservoir computing, we propose a method that uses computation results recorded by dynamic programming for combinatorial optimization problems as features for linear regression, leveraging them to assist other combinatorial optimization computations. We validate the approach on the traveling salesman and subset sum problems. Multiplexing the dynamic programming process improves approximation accuracy over generic features and reduces computation time compared with independent solutions. These results suggest a new form of computation, distinct from conventional computational design, in which multiple processes efficiently share and recycle intermediate results and states.

[LG-106] Metamorphic Testing for Clinical ML Models: A Framework Proposal and Pilot Study

链接: https://arxiv.org/abs/2607.22984
作者: Jie JW Wu,Feiyu E,Bo Chen
类目: oftware Engineering (cs.SE); Machine Learning (cs.LG)
*备注: 4 pages, 1 figure, 4 tables. Accepted at the AIware 2026 arXiv Track

点击查看摘要

Abstract:Machine learning models for clinical prediction tasks, such as in-hospital mortality and sepsis onset, routinely achieve high AUROC scores. However, AUROC measures ranking performance rather than clinical sensibility. A model may rank patients correctly overall while predicting a lower mortality risk when a patient’s SOFA score worsens, contradicting established medical knowledge. This paper proposes applying metamorphic testing (MT) to clinical machine learning models to evaluate behavioral correctness without requiring ground-truth labels for individual predictions. We design a catalog of 12 candidate metamorphic relations (MRs) for three ICU prediction tasks using the MIMIC-III and MIMIC-IV datasets, with each MR grounded in an authoritative clinical guideline. We further propose a five-layer validation strategy to ensure that MRs are clinically sound before deployment. As a feasibility study, we evaluate the approach on the UCI Heart Disease dataset. Although the three clinical models achieve strong predictive performance (AUROC = 0.849-0.900), they exhibit MT violation rates ranging from 27% to 87% across five pilot MRs. An injected-fault experiment further shows that a sign-negation error in a blood pressure feature remains undetected by AUROC but increases the MT violation rate by 31-67 percentage points. These findings suggest that metamorphic testing provides a valuable complement to conventional performance metrics for assessing the behavioral correctness of clinical prediction models. Comments: 4 pages, 1 figure, 4 tables. Accepted at the AIware 2026 arXiv Track Subjects: Software Engineering (cs.SE); Machine Learning (cs.LG) Cite as: arXiv:2607.22984 [cs.SE] (or arXiv:2607.22984v1 [cs.SE] for this version) https://doi.org/10.48550/arXiv.2607.22984 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-107] Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

链接: https://arxiv.org/abs/2607.22982
作者: Asha Barua,Sajad Khodadadian
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
*备注: 33 pages, 3 figures (each with two subfigures), 1 table. Accepted at the Reinforcement Learning Conference (RLC 2026)

点击查看摘要

Abstract:Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success. In this paper, we study exact NPG in finite-horizon Markov Decision Processes with known dynamics and horizon-dependent transition kernels. We provide the first finite-time convergence guarantees for this algorithm in this setting, for which we consider both constant and increasing step size regimes. With a constant step size \eta_t=\eta , we prove that NPG converges sublinearly with a rate of \mathcalO(H^2/t) after t iterations, where H is the horizon length. We also extend this constant step size analysis to linear MDPs in an exact population-projection oracle under a full support projection distribution, recovering the same sublinear rate as in the tabular setting. Furthermore, with increasing step sizes, we prove that this algorithm achieves a linear convergence rate of \mathcalO\left(\left(1-\frac1\vartheta_\rho\right)^t\right) for a problem-dependent constant \vartheta_\rho 1 , and the horizon-only robust schedule of the form \eta_t=\eta_0(H/(H-1))^t where \eta_00 and H \geq 2 , attains this same geometric rate.

[LG-108] Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram

链接: https://arxiv.org/abs/2607.22980
作者: Ethan Davis
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Brain-computer interfaces (BCIs) have long sought calibration-free operation, but classifiers are typically benchmarked by discrimination alone, blind to whether predicted probabilities are well calibrated - a meaningful gap given nonstationary electroencephalogram (EEG) signals and the risk of overconfident point-estimate classifiers under distribution shift. We conducted a large-scale study contrasting Bayesian complete-pooling models against frequentist baselines for cross-subject, left-hand versus right-hand motor imagery EEG classification across 20 datasets. Six frequentist pipelines were each paired with an analogous Bayesian pipeline sharing identical feature engineering, fit via Markov chain Monte Carlo posterior sampling. Our primary metric was the Brier score, decomposed into reliability and resolution, alongside AUROC for discrimination and Shannon entropy for sharpness. Each metric was analyzed via random-effects meta-analysis (REML, Knapp-Hartung adjustment), verified by leave-one-out influence analysis. Bayesian complete-pooling produced statistically but not practically significant improvements in reliability and increases in predictive uncertainty (lower sharpness); Brier score, resolution, and discrimination showed no significant differences. Between-study heterogeneity was low across all metrics, though the reliability result was sensitive to leave-one-out removal. We additionally profiled computational cost, finding that Bayesian pipelines consumed roughly thirteen times more energy than their frequentist counterparts, a cost that remains modest relative to common household appliances. These results suggest that Bayesian complete-pooling alone offers limited practical benefit for cross-subject motor imagery classification, and that partial-pooling across subjects and sessions is a more promising direction for future work.

[LG-109] Learned Interventions in Lean 4 grind

链接: https://arxiv.org/abs/2607.22972
作者: Evan Wang,Simon Chess,Sophie Szeto,Theodore Meek
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Lean~4’s \grind tactic combines congruence closure, \ematching, and case-splitting into a single automated solver, and like any such solver, it relies on hand-tuned heuristics to decide what to instantiate and where to case-split. These heuristics are tempting targets for learning, but there is a catch: because \grind’s search is non-monotone, a learned heuristic that helps one proof can break another, and an always-on replacement usually nets out near zero. We avoid this by invoking a learned intervention only after stock \grind has already failed: a failure-triggered cascade that, by construction, cannot lose a proof \grind already had. We apply it to two of \grind’s internal decisions. A cost-aware \ematch filter solves slightly more problems and runs about 5% faster. A lookahead step, proves five theorems it otherwise times out on. We also report the negative result that motivated the design: across four feature-based models, statically predicting the correct case split is no better than random, because whether a split explodes is a runtime property that the features do not capture. Our results suggest that learning within theorem-proving tactics is most effective as a mechanism for deciding when and how to spend bounded search, backed by a reliable symbolic fallback.

[LG-110] Discrepancy-Rounded Fair Bandits with Static and Time-Varying Exposure Floors

链接: https://arxiv.org/abs/2607.22935
作者: Ibne Farabi Shihab,Joyanta Jyoti Mondal,Anuj Sharma
类目: Machine Learning (cs.LG)
*备注: 28 pages, 8 figures

点击查看摘要

Abstract:Minimum-exposure constraints arise in recommendation, content curation, and regulated allocation when each provider, arm, or group must receive guaranteed exposure inside a period rather than only in aggregate. We study stochastic bandits with exact exposure floors and show that the right object is a rounding problem: a fractional fair schedule is realized as integral pulls, and the exposure error is exactly a discrepancy vector. The main contribution is a blockwise model with time-varying floors. BDQ-UCB satisfies every block floor deterministically and has fair regret governed by the nonmandatory budget R , not the horizon T , with high-probability regret O(\sqrtKR\log(KT)) . A MOSS residual variant attains O(\sqrtKR) , and a matching lower bound gives the minimax rate \Theta(\sqrtKR) , even with positive mandatory exposure; a kl-UCB ^++ residual rule adds instance-dependent optimality. The formulation becomes essential for overlapping group floors: per-arm rounding can violate a group constraint by \Omega(s) in the group size, whereas Beck–Fiala null-space rounding meets every group floor within the block budget with violation below the arm degree t , and composes with UCB at the same R -parametrized regret. For learned group plans, we close disjoint systems at \widetilde\Theta(\sqrtKT) , give a dual-ledger decomposition explaining why naive index rules fail under overlap, and prove a plan-sampling rule that is pathwise feasible under an initial cover-slack condition and attains a conditional \widetilde O(\sqrtKT) guarantee, leaving the condition-free overlap rate open. Experiments on synthetic floors, MovieLens-100k genre exposure, and deployment stress tests show exact feasibility without penalty tuning and regret competitive with tuned Lagrangian baselines.

[LG-111] Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety AAAI2027

链接: https://arxiv.org/abs/2607.22929
作者: Domenic Rosati,Ali Dadsetan,Hong Huang,Xijie Zeng,Hassan Chowdhry,Subhabrata Majumdar,Hassan Sajjad,Frank Rudzicz
类目: Machine Learning (cs.LG)
*备注: Under submission AAAI 2027

点击查看摘要

Abstract:A short fine-tuning run can undo the safety guards of an open-weight model—retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing such harmful fine-tuning while retaining benign adaptability remains difficult: the only prior method with an explicit curvature certificate, spectral deformation, inflates curvature globally and thereby obstructs benign adaptation along with harmful adaptation. We propose HarmAlign, which applies function-preserving spectral deformation along a estimated contrastive activation subspace. We derive finite-sample bounds for the estimated subspace energy and the resulting local harmful-distribution curvature lower bound. A stability–progress dichotomy for constant-step gradient descent turns the certified curvature into conditional convergence-rate control. Empirically, within a fixed-architecture, finite-budget first-order threat model, HarmAlign blocks direct fine-tuning and three data- or objective-adaptive attacks across a hazardous-knowledge relearning setting and a harmful-assistance fine-tuning setting, while the protected benign tasks remain trainable. The block persists across the tested first-order optimizer variants over every attack checkpoint, and under out-of-distribution harmful fine-tuning, and it extends to important cases in our threat model: accidental safety degradation and emergent misalignment.

[LG-112] Hidden Boundary Motion in Transformer Optimization: Function-Space Orthogonalization of Affine Weight and Bias Updates

链接: https://arxiv.org/abs/2607.22927
作者: Zhang Gongyue,Sheng Yixuan,Liu donghan,Wang Zhiyong,Ren Weihong,Liu honghai
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Weights and biases are normally optimized as separate parameter tensors, yet they do not represent separate functions when the input to an affine layer has nonzero mean. For an affine map z=Wx+b with input mean \mu , a weight update contains a sample-independent displacement \Delta W\mu that is functionally indistinguishable from a bias update. We call this hidden contribution \emphboundary motion and decompose each update into a centered, sample-varying \emphshape component and a shared \emphboundary component. On a four-layer Transformer trained from scratch on IMDb, the bias-like term g_b\mu^\top has a median norm equal to 0.664 of the raw weight-gradient norm across affine layers and training checkpoints. More strikingly, the median ratio \norm\Delta W\mu/\norm\Delta b is 134.7, while \norm\Delta W\mu/\norm\Delta b+\Delta W\mu is 0.994. Thus, under AdamW, the observed boundary motion is almost entirely realized through the weight matrix rather than the explicit bias. We implement a diagnostic optimizer, Shape–Boundary Orthogonal AdamW (SBO-AdamW), that optimizes g_W-g_b\mu^\top and g_b with independent Adam states and compensates the weight-induced boundary displacement. In a single-seed experiment, SBO-AdamW raises validation accuracy from 81.68% to 85.81% and validation-selected test accuracy from 78.73% to 82.73%, with the best validation checkpoint occurring at step 800 instead of step 3000. However, the moving-batch-center compensation produces severe bias-coordinate drift and strongly reduces boundary energy. The present evidence therefore supports hidden boundary motion as an important optimization mechanism, but it does not yet establish a final general-purpose optimizer. A stable centered-affine parameterization is identified as the required next step. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.22927 [cs.LG] (or arXiv:2607.22927v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.22927 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-113] Beyond Directed Acyclic Graphs: Causal Zeros and Causal Differential Equations

链接: https://arxiv.org/abs/2607.22910
作者: Sergei V. Kalinin
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Pearl’s structural causal model (SCM) framework, built on directed acyclic graphs (DAGs) and the do-calculus, is the dominant formal language for causal reasoning. Yet it carries two structural restrictions: every relationship must be pre-specified as a directed causal edge, and feedback cycles are forbidden. This paper examines two classes of phenomena that strain these restrictions. First, symmetric physical and economic constraints, the ideal gas law being the canonical case, carry no intrinsic causal direction. Direction emerges only under intervention, and which variable is solved for must be specified as part of the intervention. We formalize such constraints as causal zeros within an Extended Causal Model by adding an activation operator, subject to local solvability and graph-admissibility conditions. Second, for the class of finite-propagation state-space systems considered here, we treat apparent instantaneous cycles as artifacts of suppressed time and ground both causal zeros and feedback in Causal Differential Equations (CDEs). In these, the transient regime is a time-unrolled acyclic causal process, and causal zeros arise as the defining functions of attracting equilibrium manifolds; periodic and chaotic attractors define further regimes of the same dynamics, treated through attractor-relative intervention. We give the extended do-calculus, identifiability conditions, counterfactual semantics, and open problems.

[LG-114] Learning from the Descent Direction: Adaptive Gradient Descent under One-Sided Hölder Regularity

链接: https://arxiv.org/abs/2607.22906
作者: Arzu Ahmadova,Ismail Huseynov
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注:

点击查看摘要

Abstract:We study adaptive gradient descent for continuously differentiable, possibly nonconvex objectives under one-sided Hölder regularity. Unlike classical Hölder- or Lipschitz-gradient assumptions, which control the full gradient variation, our condition bounds only the directional term appearing in the descent inequality. This can allow less conservative step sizes when large gradient changes are orthogonal to, or favorable along, the update direction. We propose an adaptive scalar-step method based on an estimate of positive one-sided Hölder curvature, combined with a simple sufficient-decrease safeguard. For nonconvex objectives on a convex region containing the accepted update segments, we prove an explicit best-iterate stationarity bound with a rate determined by the Hölder exponent. Unlike predetermined diminishing step-size schemes, the method adapts to the local descent geometry. We evaluate the approach on two full-batch benchmarks designed to separate directional curvature from full gradient variation. On a binary classification problem, the method achieves the lowest final cross-entropy, objective value, and gradient norm, together with the largest classification margin among the compared scalar gradient methods. On a nonconvex Hölder regression problem, it attains the lowest final objective gap and gradient norm. These results indicate that one-sided Hölder curvature is an effective adaptive step-size signal when full-gradient variation is inflated by directions that do not hinder descent.

[LG-115] Efficient Learning of Truncated Boolean Product Distributions: Influence to the Rescue

链接: https://arxiv.org/abs/2607.22889
作者: Rohan Chauhan,Ioannis Panageas
类目: Machine Learning (cs.LG); Data Structures and Algorithms (cs.DS); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Learning the natural parameters z \in \mathbbR^n of discrete distributions \mu_z from independent samples constrained to a subset S \subseteq \0,1^n is a foundational challenge in high-dimensional statistics. Existing methods for efficiently estimating truncated Boolean product distributions, notably the work of [Fotakis et al’ COLT’20, Algorithmica ‘22], require either strong local connectivity assumptions on S – a property denoted fatness – or stringent anti-concentration assumptions and necessitate the total mass of the truncation set to be a constant with respect to n . Moreover, the results in [Fotakis et al’ COLT’20, Algorithmica '22] suffer from sample complexities that scale as \Omega(2^n) if the mass of S is exponentially small in n . In this work, we circumvent these limitations by analyzing the geometry of S under the measure \mu_z . We refine the existing parameter estimation guarantees under the fatness assumption, improving the prior sample complexity to O( \log n / \epsilon^2) for \ell_\infty -recovery, matching the untruncated minimax rate. We further generalize fatness using the notion of influence utilized in the analysis of Boolean functions and provide sufficient conditions for efficient inference. Notably, unlike previous work, our method does not require sampling at arbitrary parameterizations of the model. Lastly, we establish a theoretical lower bound demonstrating the sample complexity exhibits an intrinsic exponential dependence on the width of the model and the minimum distance between elements in the set. Subjects: Machine Learning (cs.LG); Data Structures and Algorithms (cs.DS); Machine Learning (stat.ML) Cite as: arXiv:2607.22889 [cs.LG] (or arXiv:2607.22889v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.22889 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-116] Queryable Self-Organizing Maps: A Database Abstraction for Topology-Driven Data Exploration

链接: https://arxiv.org/abs/2607.22843
作者: Denis Mayr Lima Martins,Gottfried Vossen
类目: Databases (cs.DB); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Self-Organizing Maps (SOMs) have long been used as exploratory tools for high-dimensional data: they organize objects into a two-dimensional topology that reveals clusters, gradients, sparse regions, dense regions, and boundaries. Yet, in modern data systems, SOMs are typically trained and visualized outside the DBMS, disconnected from the relational data they summarize. We introduce the abstraction of a queryable data map: a learned topological artifact consisting of representatives, neighborhood relations, object assignments, and derived summaries. We instantiate this idea with MapDB, a lightweight prototype that makes SOM artifacts queryable so users can explore data topology without leaving the database. Experimental study shows that SOM training is feasible at moderate analytical scale, that map queries are interactive after materialization, and that SOM regions provide meaningful targets for exploratory SQL.

[LG-117] MEMENTO: Memory-Guided Memetic Code-as-Policy Evolution

链接: https://arxiv.org/abs/2607.22832
作者: Alkis Sygkounas,Victor Aregbede,Amy Loutfi,Andreas Persson
类目: Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Long-horizon embodied tasks require policies that execute many dependent actions before task success can be observed. Representing policies as executable control pro- grams (code-as-policy) enables their decision logic to be inspected and revised after rollout evaluation. Revised programs can then be executed and compared by rollout performance, framing policy improvement as execution-guided program search. Evo- lutionary methods driven by large language models (LLMs) provide a natural mecha- nism for this search by generating variants and selecting high-performing candidates. However, existing approaches primarily select among independently generated vari- ants and lack a sequential local improvement phase. We introduce MEMENTO, a memory-guided single-elite memetic framework for code-as-policy evolution. ME- MENTO first evolves a rollout evaluator that maps policy rollouts to scalar fitness and structured feedback metrics. Fitness selects accepted candidates and the next elite, while feedback metrics condition policy proposals generated by memory-guided hill-climbing, macro-mutation, and crossover. We evaluate MEMENTO on two long- horizon embodied domains: Robosuite Franka Tower-of-Hanoi manipulation and AI2- THOR household interaction. MEMENTO outperforms Eureka and REvolve, adapted as code-as-policy evolutionary baselines, in task success and generalization to held- out Robosuite object configurations and unseen AI2-THOR scenes. Ablations show that zero-shot generation and unevolved evaluators fail to solve either domain, and that removing policy-search branches reduces performance. Finally, we deploy the best-evolved Robosuite policy on a physical Franka robot, demonstrating the feasibil- ity of sim-to-real transfer of the evolved code-as-policy. Code, prompts, and videos are available at: this https URL.

[LG-118] SLA-Constrained Carbon-Aware Routing in Geo-Distributed Serverless Clouds

链接: https://arxiv.org/abs/2607.22806
作者: Anmol Chaudhary(Department of Electronics and Computer Engineering, NIAMT Ranchi),Rahul Mishra(Department of Electronics and Computer Engineering, NIAMT Ranchi)
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
*备注:

点击查看摘要

Abstract:Modern cloud deployments distribute applications across multiple geographic regions, yet standard routing mechanisms prioritize latency while ignoring the fluctuating carbon intensity of local power grids. Latency-driven routing incurs avoidable carbon emissions, particularly when cleaner regions are within acceptable latency bounds. The proposed model formulates the carbon-aware serverless routing problem as a constrained optimization over geo-distributed cloud regions and introduces an SLA-constrained carbon-aware routing policy that achieves optimal carbon reduction within the SLA-feasible region, evaluated using real carbon intensity measurements across 5 primary AWS deployments. Experimental results show that the proposed policy achieves up to 46.8% carbon reduction while maintaining zero SLA violations across all evaluated thresholds. The system reduces carbon by an average of 27.4% under mixed workloads, and the routing overhead is very low (less than 0.02% of total request latency). A scalability study across 12 AWS regions spanning 6 continents demonstrates that average carbon savings increase from 27.4% to 47.5% as routing flexibility expands under mixed workloads. The proposed work contributes to SDG 13 (Climate Action) and SDG 7 (Affordable and Clean Energy) by enabling low-carbon routing decisions. These results indicate that cloud systems can achieve significant carbon savings without compromising user experience.

[LG-119] FusionML: Prefill Not Decode - Mechanism and Boundaries of CPUGPU Co-Execution on Unified-Memory Apple Silicon

链接: https://arxiv.org/abs/2607.22785
作者: Om Mohite
类目: Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Performance (cs.PF)
*备注:

点击查看摘要

Abstract:Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units. Prior attempts, including our own, failed or produced precision-confounded wins. We identify the cause: MLX’s lazy-graph scheduler \emphserializes cross-stream work whenever a CPU-stream operation consumes an unmaterialized GPU result inside one evaluation graph, so a row-split matmul that runs \x1.38 faster with materialized inputs runs \x0.66 slower than GPU-only inside a lazy graph; an eager materialization boundary restores concurrency (\x1.34). \sys implements a per-layer, contention-aware CPU+GPU row split for transformer prefill built on this fix. Evaluated across five chips and three Apple-Silicon generations, community-replicated, the split accelerates Llama-shaped decoder-block prefill by \x1.15–\x1.38, unchanged at full 32-block depth, and reaches \x1.18–\x1.25 faster time-to-first-token on a real Qwen2.5-7B checkpoint served through stock MLX-LM, with token-identical outputs and unchanged decode throughput. We characterize the boundaries equally carefully: decode cannot benefit, bound by shared bandwidth co-execution does not add; precision-matched training loses \x0.86–\x0.97 on all five chips; ANE dispatch overhead excludes it at layer granularity; and a no-regression runtime gate becomes self-defeating under memory pressure, where probing an alternative mode evicts the active mode’s working set. Code, raw results, and generation transcripts are released.

[LG-120] Predicting the Outcome of rTMS Depression Therapy using EEG Signals and CNN

链接: https://arxiv.org/abs/2607.22776
作者: Wael Korani,Md Fahimul Kabir Chowdhury,Sadam AlQadi,Priyan Malarvizhi kumar,Reza Rostami,Reza Kazemi
类目: Machine Learning (cs.LG)
*备注: Presented at 8th International Conference on Recent Trends in Image Processing Pattern Recognition (RTIP2R)

点击查看摘要

Abstract:Repetitive transcranial magnetic stimulation (rTMS) is a non invasive therapy for Major Depressive Disorder (MDD). In this study, we generate images using two time frequency methods to represent EEG signals: Fourier-Bessel Series Expansion with Euclidean Distance (FBSE-ED) and Discrete Wavelet Transform (DWT). We propose an efficient deep learning classifier to predict the outcome of rTMS depression therapy. In this study, we use a private rTMS databases to train a lightweight custom Convolutional Neural Network (CNN) using 10-fold cross validation strategy in order to avoid any bias in our results. The results show that the FBSE-ED representation achieves the highest classification accuracy of 93.60%, outperforming traditional time-frequency technique (DWT). In addition, the proposed architecture with FBSE-ED image representation technique outperforms more complex EEG-Specific deep learning models (EEGNet, DeepConvNet, SleepEEGNet) by 3.62-10.72% and pretrained models (Xception, DenseNet201, and MobileNetV2) by 23.03-27.35%. For more experiments, we utilize another private rTMS database as test database to show the robustness of the proposed model. Our results suggest that integrating advanced signal decomposition with deep learning can facilitate early prediction of rTMS treatment response and support more targeted clinical decision-making. The proposed framework is interpretable, computationally efficient, and well-suited for deployment in real-world local psychiatric clinics.

[LG-121] CC-AOS: Cost- and Horizon-Conditioned Amortized Backward Induction for Finite-Horizon Optimal Stopping

链接: https://arxiv.org/abs/2607.22774
作者: Tianwei Yu
类目: Machine Learning (cs.LG)
*备注: 17 pages, 2 figures, and 4 tables

点击查看摘要

Abstract:Finite-horizon optimal stopping is a central problem in early time-series classification, where a system must decide at each sequence prefix whether the expected benefit of another observation justifies its acquisition cost. Existing data-driven backward-induction methods typically solve each cost-horizon operating point separately, so changing operating conditions requires repeated optimization and separate model stacks, making continuous cost adaptation and multi-horizon deployment inefficient. We propose CC-AOS (Cost- and Horizon-Conditioned Amortized Optimal Stopping), a structured amortized solver for a family of finite-horizon stopping problems with continuous costs and multiple horizons. CC-AOS learns a shared continuation-value model conditioned on the current state, absolute time, remaining horizon, and acquisition cost through joint amortized fitted backward induction. We establish that the exact value and continuation functions are nondecreasing, concave, and horizon-dependently Lipschitz in cost, encode these properties in the model architecture, and derive residual-based bounds on value and policy errors. Experiments on controlled Gaussian and time-varying non-Gaussian processes and the FordA engine-noise time-series benchmark compare CC-AOS with representative per-operating-point backward-induction solvers and tuned static stopping rules. At six unseen FordA cost-horizon pairs, one CC-AOS checkpoint achieved a lower terminal-risk-plus-sampling-cost objective than independently fitted Convex Function Learning at all six pairs, with an average reduction of 15.75 percent, while matching the tuned static thresholds on average.

[LG-122] An Integrated Deep Learning and Statistical Framework for Whole-Network Gene–Environment Association with Leaf Vascular Architecture

链接: https://arxiv.org/abs/2607.22763
作者: Geran Zhao,Yangsheng Wang,Xiaotian Dai,Guifang Fu
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Leaf veins exhibit remarkable diversity in architecture and patterning, yet existing gene–environment association studies have primarily quantified leaf venation using a small collection of low-dimensional summary traits, thereby discarding most of the structural information contained in the original images. We propose an integrated deep learning and statistical framework. The proposed framework achieves four methodological advances. First, it represents the complete leaf vascular architecture as a whole-network image phenotype. Second, it fine-tunes the deep learning-based Edge Detection with Transformers (EDTER) model to accurately extract whole-network leaf vascular architecture from RGB images by jointly learning local and global contextual features. Third, it constructs a new annotated leaf image database by integrating edge maps generated by DiffusionEdge with the Berkeley Segmentation Database (BSDS500). Fourth, it applies Semiparametric Sparse Canonical Correlation Analysis (SSCCA) to perform variable selection and model associations between repeatedly measured high-dimensional Bivariate image responses and high-dimensional predictors while simultaneously accommodating sparse, zero-inflated data represented by edge maps through a truncated latent Gaussian copula model. Two simulation studies demonstrate the performance of the proposed framework under increasing levels of complexity. Application to a real \emphPopulus dataset identifies three significant gene–geography interactions associated with leaf vascular architecture, providing new biological insights and establishing a broadly applicable methodological framework for high-dimensional complex image phenotypes.

[LG-123] DRC-Aid: Design-Rule Correction via Agent ic Framework utilizing Inference-Time Large Language Models

链接: https://arxiv.org/abs/2607.22761
作者: Anushka Mukherjee,Kang He,Kaushik Roy
类目: Hardware Architecture (cs.AR); Machine Learning (cs.LG)
*备注: 7 pages

点击查看摘要

Abstract:Resolving Design Rule Violations (DRVs) in layouts entails an iterative loop of geometric edits and verification. We present DRC-Aid, a closed-loop agentic framework that automates local DRC repair by formulating it as verification-in-the-loop search. To constrain the combinatorial geometric repair space, a deterministic Rule Engine converts physical verification tool-reported violations into a bounded menu of geometric edits. An off-the-shelf Large Language Model (LLM) evaluates local geometric context to select edits from this menu, with budgeted depth-first search and backtracking. Immediate feedback from verification tools such as Calibre nmDRC/nmLVS enforces geometric compliance and guards against electrical-topology degradation, while a global Memory Bank prevents cyclic re-exploration. Evaluated on FreePDK45 layouts containing DRVs, DRC-Aid achieves DRC-clean, LVS-equivalent repairs in ~92.5% of cases with a ~98% total violation reduction, while residual cases yield partially repaired LVS-equivalent candidates. Under an identical search and verification infrastructure, LLM-based selection outperforms random (54.4%) and deterministic-heuristic (83.3%) policies, with the gap widening on cases with six or more violations.

[LG-124] Benchmarking LLM s for Verilog Design Flows

链接: https://arxiv.org/abs/2607.22759
作者: Angshuman Chakravertty,Rahul Koshti,Buddhi Prakash Sharma,Vinay Chamola
类目: Hardware Architecture (cs.AR); Machine Learning (cs.LG)
*备注: 7 pages, 3 figures, 3 tables. Manuscript prepared for submission to IEEE Design Test

点击查看摘要

Abstract:Large language models (LLMs) show promise in code generation, but their capabilities to produce correct, synthesizable hardware description language (HDL) code still remain to be properly benchmarked. Existing evaluations are primarily relying on pass@k metrics and lack proper end-to-end toolchain validation. This paper presents a reproducible benchmarking platform that evaluates open-source LLMs on Verilog RTL generation across 50 curated tasks consisting of combinational, sequential, finite state machine (FSM), and mixed designs. The pipeline consisting of constrained prompting, post-processing, and semantic-aware iterative refinement with waveform analysis, formal equivalence verification, and Abstract Syntax Tree (AST)-based repair validates the generated code via Verilator compilation and Icarus Verilog simulation. Across the 12 benchmarks and the 1,610 total runs evaluating three models of different sizes (Llama-3-8B, StarCoder2-7B, and TinyLlama-1.1B), the pipeline improved syntax validity from 0% to a 70.43% average and simulation pass rate to 51.8% across three open-source models. Most notably TinyLlama (1.1B parameters) achieved the highest individual syntax validity at 80.0%, with functional correctness comparable to the 8B model. The platform and dataset are open-source, enabling reproducible evaluation of generative AI for hardware design workflows.

[LG-125] Stacking the Deck: Tunable Trainability in Stacked LCUs

链接: https://arxiv.org/abs/2607.24686
作者: Nikhil Khatri,Stefan Zohren,Gabriel Matos
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注: 19 pages

点击查看摘要

Abstract:Variational quantum circuits have been central to many proposed near-term applications of quantum computing, but a growing body of evidence suggests that trainability and quantum advantage are fundamentally at odds: ansätze expressive enough to resist efficient classical simulation tend to exhibit barren plateaus, while structures that provably rule out barren plateaus typically render them classically simulable. We propose a stacked linear combination of unitaries (S-LCU) as a variational ansatz which provides a tunable trade-off between barren plateaus and classical simulability. Using a diagrammatic analysis, we bound the loss-landscape variance of the Free Fermion S-LCU, whose elements are fermionic Gaussian unitaries. We prove a variance lower bound of \Omega(1/(n k^3l)) , with a simulation cost of O(k^2l n^3) using the best known classical algorithm, compared to a quantum gate complexity of only O(lkn^2) . The number of layers l serves as a single dial that trades computational complexity against the rate of cost concentration. This offers practitioners a systematic method for constructing ansätze with a complexity-trainability trade-off that best suits their application and hardware.

[LG-126] A Model for Imbalanced Label Aggregation: A Focus on Minority-Class Detection

链接: https://arxiv.org/abs/2607.24622
作者: Gabriel Singer,Samuel Gruffaz,Olivier Vo Van,Nicolas Vayatis,Argyris Kalogeratos
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study imbalanced crowdsourcing with a focus on class-dependent annotator accuracy, a setting that, to the best of our knowledge, remains relatively underexplored despite its importance in real-world inspection systems where the labels of greatest operational importance are also the rarest ones. In this setting, annotators may be reliable on both classes, unreliable on both classes, majority-class specialists, or minority-class specialists. Existing models only partially address this problem: they either capture class-dependent errors but ignore item difficulty, or they model item difficulty without capturing class-dependent errors. To fill this gap for imbalanced datasets in crowdsourcing, we introduce a generative aggregation model combining item difficulty with class-dependent annotator competence. The model allows both annotator abilities and item difficulties to vary across classes. We then revisit Condorcet’s Jury Theorem in the class-imbalanced setting. We also show that majority voting asymptotically preserves the underlying class proportion. We evaluate our model on 33 real-world crowdsourcing datasets, covering multiclass tasks such as images and text, as well as two large-scale regimes: large-scale annotation datasets, with many annotations per item, and large-scale item datasets, with a large number of annotated instances. Across these diverse settings, our model consistently achieves the highest minority recall while remaining competitive in balanced accuracy, making it particularly relevant when rare-label recovery is the primary objective.

[LG-127] he balance between compactness and forecast accuracy of data-driven latent-space reduced-order models in controlled wake flows

链接: https://arxiv.org/abs/2607.24569
作者: Alberto Solera-Rico,Patricia García-Caspueñas,Carlos Sanmiguel Vila,Stefano Discetti
类目: Fluid Dynamics (physics.flu-dyn); Machine Learning (cs.LG)
*备注: 32 pages, 20 figures

点击查看摘要

Abstract:Model-based active flow control requires predictive models that are accurate, stable, and fast enough for real-time optimisation. In controlled wake flows, this is often achieved through Reduced-Order Models (ROMs) that first compress high-dimensional velocity snapshots into a latent space and then learn a time- stepping predictor for the dynamics in the latent space. Here, we study how the choice of the spatial encoder affects the predictability of the resulting latent coordinates for wake flows under control inputs. Using two actuated 2D wake configurations, a simplified truck wake and the fluidic pinball, we compare Proper Orthogonal Decomposition (POD) against nonlinear Convolutional Autoencoders (CAEs) and two types of variational autoencoders for compression, and evaluate several temporal predictors based on Long Short-Term Memory networks. CAEs achieve higher compression efficiency and sharper short-term reconstructions, but they produce latent dynamics that are more irregular and with broadband spectral content. As a consequence, long-horizon forecasts degrade faster and show a higher probability of catastrophic divergence than POD-based models. POD yields smoother latent trajectories that are easier to learn and extrapolate, leading to more reliable predictions beyond the short- term regime. These results reveal a clear trade-off between compactness and forecast accuracy, and suggest that the stability of the latent dynamics prediction can outweigh maximal compression. This is particularly relevant for control strategies rooted in forecasts of the dynamics, such as model predictive control and reinforcement learning. The findings provide practical guidance for designing actuation-aware, hardware-feasible predictive ROMs for real-time flow control.

[LG-128] Self-Attention Dynamics with Rotary Position Embeddings: Twisted States and Explicit Consensus Rates on the Sphere

链接: https://arxiv.org/abs/2607.24502
作者: Hao Ye(Xi’an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences, University of Chinese Academy of Sciences)
类目: Dynamical Systems (math.DS); Machine Learning (cs.LG)
*备注: 28 pages, 4 figures. Includes complete proofs and independent numerical cross-checks

点击查看摘要

Abstract:Rotary position embeddings (RoPE) modify attention scores through position-dependent rotations, but their effect on normalized token dynamics is not captured by the vanilla spherical self-attention model. We study the continuous-time dynamics obtained when queries and keys are rotated while values remain on the unit sphere. The resulting attention kernel is reversible and admits a sharp uniform softmax floor, yet the natural RoPE interaction energy has derivatives of both signs within one fixed nontrivial system. Every consensus state remains an equilibrium, and its transverse linearization is a reversible Markov operator whose kernel depends on the consensus point through its energy across RoPE planes. On a resonant single-frequency ring we derive an exact Bessel-aliasing spectrum, including non-coprime frequencies and the correct fixed-ring large- \beta asymptotics. Globally, closed hemispheres are invariant, while pairwise non-obtuse configurations and strict open semicircles contract with explicit half-angle and single-point tail bounds. These regional estimates instantiate a kernel-generic positivity principle with the sharp RoPE softmax floor. RoPE also selects an explicit score-flattening twisted branch; the generic resonant family is non-hyperbolic and linearly unstable, whereas an odd antipodal family becomes a hyperbolic saddle after quotienting global rotation. In multiple dimensions, the local consensus gap can depend non-monotonically on the allocation of energy across frequency planes, so no universal ordering by frequency is valid. Independent matrix, finite-difference, and nonlinear-flow computations cross-check the theorem boundaries and the reported constants.

[LG-129] Frequency-Based Reservoir computing

链接: https://arxiv.org/abs/2607.24420
作者: Arthur S Powanwe
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 24 pages, 8 figures

点击查看摘要

Abstract:Reservoir computing has emerged as an efficient machine learning framework for predicting time series generated by dynamical systems. In contrast to other machine and deep learning approaches, a reservoir computing trains only the output layer via linear regression, leaving the reservoir (recurrent layer) untrained. This simplification makes reservoir computers easier to train and more amenable to experimentation. However, because current reservoirs consist of networks of randomly connected nodes and require the optimization of numerous hyperparameters, a framework that precisely explains how reservoir computing operates and how it can be optimized remains missing. Here, we propose a frequency-based reservoir inspired by the brain’s oscillatory dynamics and its hierarchy of timescales. The frequency-based reservoir can be interpreted as an ensemble of independent oscillatory units, each processing a portion of the input’s frequency content. This allows us to understand the reservoir’s internal behavior by modeling it as a single unit driven by an external input. Borrowing from the theory of a nonlinear oscillator forced by complex periodic inputs, we found that units of the frequency-based reservoir selectively amplify and store specific input frequencies, which are then used for prediction. The frequency-based reservoir performs as well as or better than equivalent random reservoirs. Furthermore, the frequency-based approach can be optimized to improve short-term prediction, a property that random reservoirs lack. Finally, we show that the frequency-based reservoir can also predict complex spatiotemporal dynamics. Our results show that reservoir computing can be designed using brain properties and theoretical insights borrowed from the physics of forced nonlinear oscillators. Comments: 24 pages, 8 figures Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG) Cite as: arXiv:2607.24420 [stat.ML] (or arXiv:2607.24420v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2607.24420 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-130] proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference

链接: https://arxiv.org/abs/2607.24401
作者: Alexandra N. M. Darmon,Deeksha Sinha,Steve Wilkins-Reeves,Caner Gocmen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 15 pages, 2 figures

点击查看摘要

Abstract:Proxy outcomes (such as short-term behavioral signals, model predictions, or surrogate endpoints) are frequently used in place of primary outcomes that are too slow to mature, rare, or challenging to measure directly. But valid inference on a proxy does not guarantee valid inference on the primary estimate as proxy-based estimates can be systematically biased in ways that are difficult to predict, leading to improperly calibrated confidence intervals. We present proxymate, a framework and open-source Python package for proxy validation and adjustment. proxymate organizes into four levels: The Representativity Level (population validity), the Unit Level (measurement quality), the Estimate Level (decision validity), and the Domain Level (cross-domain transportability). Within each level, proxymate provides diagnostic checks, and targeted adjustment strategies that map specific failures to appropriate corrections. At Meta, proxymate has been adopted by many different use cases, spanning experimentation, prevalence estimation, and monitoring use cases, all facing different proxy challenges (limited human review time, long maturation window of outcomes, low detectability) and showcasing the modularity of the framework. Across all products, proxymate assessed and corrected millions of proxy, primary unit comparisons. It has facilitated launches across multiple work streams including enabling quick decision making on thousands of experiments. Comments: 15 pages, 2 figures Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG) Cite as: arXiv:2607.24401 [stat.ML] (or arXiv:2607.24401v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2607.24401 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-131] Stochastic Counterdiabatic Driving via Biorthogonal Liouvillian Eigenmodes

链接: https://arxiv.org/abs/2607.24393
作者: Sandeep Suresh Cranganore,Sebastian Lehner,Johannes Brandstetter,Max Welling
类目: Computational Physics (physics.comp-ph); Statistical Mechanics (cond-mat.stat-mech); Machine Learning (cs.LG)
*备注: 34 pages, 13 figures, 14 Tables (including Supplimentary Material)

点击查看摘要

Abstract:Finite-time driving of stochastic systems generates excess dissipation, causing the evolving probability distribution to lag behind the instantaneous equilibrium, and consequently degrading the convergence of nonequilibrium free energy estimators based on the Jarzynski equality. Escorted free energy simulations address the non-adiabatic lag by engineering control fields \mathbfu that eliminate the lag, enforcing the trajectory-wise equality \mathcalW_\mathbfu = \Delta \mathcalF , and yielding zero-variance estimators. However, constructing the escorting field in closed form remains a challenge, approached variously through flow-field methods, targeted free energy perturbation, or learned diffeomorphisms. In this work, we construct a complementary numerical framework based on gauge-type transforms instead of generalized coordinate transforms for perfect escorting based on the exact spectral decomposition of the time-dependent Fokker-Planck generator. The biorthogonal decomposition of the Liouville operator directly yields a counterdiabatic correction whose action on the instantaneous equilibrium distribution exactly cancels the non-adiabatic lag at arbitrary driving speed in formal analogy with shortcuts-to-adiabaticity techniques such as Berry’s transitionless driving for quantum systems. Numerical verification for simulations of an overdamped particle in a time-varying double-well potential and harmonic traps confirms that the counterdiabatic condition is satisfied to machine precision, with the non-adiabatic lag suppressed by roughly twelve orders of magnitude in total variation distance and sixteen orders in KL divergence relative to the unescorted dynamics. As a diagnostic, we demonstrate vanishing dissipated work \mathcalW_\textdiss(t) \approx 0 for the deterministically propagated Fokker-Planck density across all protocol speeds.

[LG-132] Catalyst Diffusion Transformer: Generative Inverse Design of Heterogeneous Catalysts

链接: https://arxiv.org/abs/2607.24272
作者: Hayoung Doo,Dong Hyeon Mok,Seoin Back,Jonggeol Na
类目: Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:The vast chemical design space and complex, interdependent design variables make catalyst discovery for targeted properties highly labor- and resource-intensive. Although generative models have emerged as a promising solution, existing approaches are generally limited to single-property conditioning or narrow chemical spaces. Here, we present Catalyst Diffusion Transformer (CatDiT), a unified framework for inverse catalyst design that generates valid and novel structures ranging from intermetallic alloys to oxide surfaces. By learning compressed latent representations, CatDiT enables efficient training and rapid sampling while supporting simultaneous conditioning on adsorbate type, binding energy, and catalyst class. The model provides reliable control of discrete properties and directional control of continuous properties, enriching candidate pools for reaction-specific catalyst discovery. As a representative application, multi-conditional generation for the nitrogen reduction reaction (NRR) yields 28 density functional theory (DFT)-relaxed alloy candidates that satisfy the target activity window and lie above the pure-metal *N-*H scaling line, corresponding to a ~1.5-fold enrichment over the source distribution. These results establish CatDiT as a practical and scalable approach for property-directed catalyst inverse design and targeted catalyst generation.

[LG-133] Decision trees Frobenius traces and Weierstrass coefficients of elliptic curves

链接: https://arxiv.org/abs/2607.24251
作者: Barinder S. Banwait,Xiaoyu Huang,Kyu-Hwan Lee,Seewoo Lee,Thomas Oliver,Alexey Pozdnyakov
类目: Number Theory (math.NT); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We investigate the extent to which the reduced minimal Weierstrass coefficients of an elliptic curve over \mathbbQ may be computed from it’s Frobenius traces. Decision tree models reveal that the first two reduced minimal Weierstrass coefficients can be recovered with perfect accuracy from the Frobenius traces at the primes 2 and 3 , and the third by supplementing these two traces with the conductor parity. We subsequently prove explicit formulae for these coefficients using the Frobenius traces and conductor parity. These formulae appear to be new. In particular, we deduce that the first three reduced minimal Weierstrass coefficients of an elliptic curve are determined by its isogeny class.

[LG-134] Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD HSIC KSD

链接: https://arxiv.org/abs/2607.24235
作者: Jose Cribeiro-Ramallo,Florian Kalinke,Zoltán Szabó
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:

点击查看摘要

Abstract:Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others. Their fastest estimators are known to converge at a parametric rate— n^-1/2 —under mild conditions. While this rate is known to be minimax optimal on \mathbb R^d under strict assumptions with bounded kernels, little is known about its optimality beyond the finite-dimensional Euclidean setting with unbounded kernels. In this work, we prove that the minimax lower bound of estimation of the most popular kernel discrepancies (maximum mean discrepancy, Hilbert-Schmidt independence criterion and kernel Stein discrepancy; MMD, HSIC, KSD) is n^-1/2 on general topological spaces, and under mild assumptions on the kernel; the same rates are shown (as corollaries) to hold for the estimation of the mean embedding and the centered cross-covariance operator. Our results settle the question of optimal estimation of these kernel discrepancies.

[LG-135] On Non-Stationary Dynamic Pricing: Adaptivity and Optimality

链接: https://arxiv.org/abs/2607.24115
作者: Feiyu Jiang,Zifeng Zhao
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We study the contextual dynamic pricing problem under non-stationarity, where a firm sells products to T sequentially arriving consumers that behave according to an unknown demand model that can change over time. The demand model is assumed to be a generalized linear model (GLM), allowing for a feature vector in \mathbbR^d that encodes products and consumer information. To achieve optimal revenue (i.e., least regret), the firm needs to learn and exploit the unknown GLMs while monitoring for potential changes. We propose a multiscale change-point detection based algorithm that achieves a regret of order \widetildeO(\sqrts_TdT\wedge\V_T^1/3d^1/3T^2/3+\sqrtdT) , where s_T is the number of piecewise stationary segments and V_T is a newly defined notion of design-adjusted variation budget of model parameters. Our algorithm is adaptive and does not require knowing s_T or V_T . Moreover, to our knowledge, this is the first dynamic pricing algorithm that is adaptive to the nature of changes and achieves the best-of-both-worlds rate, thus closing a long-standing gap in the literature. We remark that, due to the varying contexts, existing works in the adaptive non-stationary bandit literature cannot be applied to achieve optimality for contextual dynamic pricing. The regret is further accompanied with a newly constructed minimax lower bound, confirming the optimality of our algorithm (up to logarithmic factors). Extensive numerical experiments are conducted to illustrate the efficiency and robustness of the proposed algorithm in non-stationary dynamic pricing.

[LG-136] Variational Quantum Conditional Boltzmann Machines for Time-Series Forecasting: Architectures Symmetric Hyperparameter Evaluation and a Nonlinear Benchmark

链接: https://arxiv.org/abs/2607.24065
作者: Gerhard Hellstern,Danyal Maheshwari,Martin Zaefferer,Martin Braun,Tanja Döhler
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG); Statistical Finance (q-fin.ST)
*备注: 45 pages

点击查看摘要

Abstract:In this study, we developed and evaluated four conditional energy-based forecasting architectures: a classical Gaussian-Bernoulli CRBM, a hybrid quantum-classical QCRBM, a full-register QQRBM, and a lag-feature QFeatureQRBM with complete derivations of their conditional distributions, Contrastive-Divergence gradients, and hybrid training, bridging the energy-based formulation and the implementation-level quantum computation. Unlike prior comparisons, our evaluation enforces symmetric hyperparameter optimisation: classical and quantum-specific hyperparameters receive an equally thorough grid search across thirteen structured experiments. We test on two data classes, a Gaussian-process dataset (GP) generated with real financial data and the input-driven NARMA-10 nonlinear benchmark. Across both regimes we find no systematic evidence of a quantum advantage at the available sample size: no quantum architecture improves on the best classical baseline. The fully quantum QQRBM and QFeatureQRBM are significantly worse, whereas the hybrid QCRBM is statistically indistinguishable from the strongest classical CRBM on both datasets. A power analysis bounds this null result: at n = 12 only medium-to-large effects are detectable, so small advantages cannot be excluded. An iso-parameter (matched-budget) comparison reaches the same conclusion: the classical CRBM is lowest at three of the four budgets and no CRBM-vs-QCRBM difference is significant at any budget.

[LG-137] he Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression

链接: https://arxiv.org/abs/2607.24041
作者: Kevin Han Huang,Haoyu Ye,Somak Laha,Morgane Austern
类目: atistics Theory (math.ST); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate. In this paper, we relax both assumptions and derive deterministic equivalents for the prediction risk in a vanishing-ridge regime. We show that degeneracy of the covariance matrices and dependence can lead to multiple descent, and characterize where the corresponding peaks can occur. Our proofs use a novel graph representation of the variance profile. We show that maximum matchings and the Dulmage–Mendelsohn decomposition of the associated bipartite graph identify the configurations at which the variance becomes singular.

[LG-138] Smooth Learning with Hard Constraints via Legendre-Regularized Policies

链接: https://arxiv.org/abs/2607.24007
作者: Zikun Lin,Rui Chen,Yijie Wang
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We revisit contextual optimization from the perspective of policy class design. A desirable policy class should be expressive enough to learn rich context-decision relationships, should enforce hard feasibility constraints rather than soft penalty terms, and should remain smooth enough for gradient-based training on downstream decision losses. Existing approaches usually emphasize only part of these requirements. We propose Legendre-regularized policies, which parameterize decisions as solutions of regularized optimization problems over the original feasible region. This construction yields policies that are feasible by construction and differentiable with respect to learned latent parameters. We prove that the associated optimizer map is single-valued, maps onto the relative interior of the feasible set, admits an explicit Jacobian, is Lipschitz continuous, and can be made arbitrarily smooth. We also establish a universal approximation result showing that the proposed class can approximate any continuous feasible policy on compact context sets. The framework unifies explicitly regularized optimizers and implicit perturbation-based smooth optimizers. Experiments on contextual newsvendor and resource allocation problems show that our approach improves prescriptive performance relative to the benchmark methods.

[LG-139] HydroAgent : Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows

链接: https://arxiv.org/abs/2607.23983
作者: Qingyi Yang,Siqian Qiu,Bing Li,Xu Shan,Jia Feng,Shunan Zhou,Xudong Zhou,Tiantian Xing,Jiale Guo,Xiaoyi Dong,Gaoyu Liu,Xiaohuan Liu,Haiqing Pu,Qingwen Deng,Xun Zhang,Zhongrun Xiang,Haiyang Qian,Ying Yan,Yongkang Xu,Nuo Lei,Tianlong Jia,Baoying Shan,Carlo De Michele
类目: Geophysics (physics.geo-ph); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer. Although artificial intelligence methods have advanced flood prediction and model-error correction, most existing studies have not explicitly represented the tacit expert rules, review checkpoints, and workflow constraints that connect model outputs to operational warning decisions. To address this issue, we propose HydroAgent, a skill-orchestrated agent framework that embeds Large Language Models (LLMs) into a model-driven flood forecasting workflow, where each skill encodes explicit rules to bound LLM reasoning. We validated its effectiveness using five state-of-the-art LLMs in the South Yamhill River basin. Our results demonstrate that prior judgment captures observed peak flow and flood volume within 5% tolerance in 10 and 11 out of 14 events, with 5-fold cross-validation over 129 events yielding Pearson correlations of 0.62 and 0.84. Building on a high-baseline scheme library (average KGE 0.890), the guided scheme selection further improves KGE by 0.023-0.154, with simulated peak flow and flood volume falling within the prior judgment ranges for 14 and 13 out of 14 events. All five tested LLMs successfully execute the HydroAgent workflow with comparable judgment accuracy (40%-80%), while showing moderate performance variation and substantial cost differences. HydroAgent does not aim to replace human forecasters; instead, it translates their tacit expertise into an auditable and reproducible workflow, streamlining analytical steps and supporting more informed decision-making. This skill-orchestrated paradigm demonstrates how explicit rule boundaries can guide language model reasoning to complement physically based simulation in next-generation flood forecasting.

[LG-140] Distributional Split Criteria for Random Forests: Extensions Shrinkage and the Robustness of Mean Splitting

链接: https://arxiv.org/abs/2607.23721
作者: Silas Koemen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 1 figure

点击查看摘要

Abstract:Distributional random forests replace mean-based CART splitting with criteria that compare the full conditional response distribution in candidate children. We implement and systematically study a family of such criteria inside a single honest-forest implementation: isotropic random-Fourier-feature maximum mean discrepancy (MMD), an anisotropic diagonal-bandwidth variant, an adaptive per-split frequency-selection variant, and a non-kernel sliced-Wasserstein criterion, together with post-hoc kernel-mean shrinkage of the forest weights. Using paired-seed comparisons across synthetic quantile mechanisms, real univariate benchmarks, a California-housing subsample curve, and multivariate synthetic and real responses, we characterize where each extension pays. Three findings recur. First, among distributional criteria ordinary isotropic MMD is already close to best in class: the anisotropic, adaptive-frequency, and sliced-Wasserstein extensions, and post-hoc shrinkage, do not systematically improve on it. Second, on scalar tabular regression mean-based CART splitting remains the robust default and wins many cells. Third, multivariate responses are the regime where distributional splitting clearly earns its keep, most sharply on a pure-dependence copula where the energy score separates the criteria even though marginal CRPS does not. The evidence supports a simple allocation story: distributional splitting helps only when non-location structure is both present and estimable; otherwise it dilutes split-selection power away from the mean. All criteria, the honest forest, and the paired-comparison harness are implemented in the open-source \textttdrforest library, whose Rust-backed split search makes broad criterion sweeps inexpensive.

[LG-141] When Rates Are Geometric: Rate-Certificate Transfer for Contact Splittings in Optimization

链接: https://arxiv.org/abs/2607.23642
作者: George A Kevrekidis
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Dynamical Systems (math.DS)
*备注:

点击查看摘要

Abstract:Discrete optimization algorithms are often analyzed through continuous-time limiting ODEs, but a convergence certificate for the ODE is not automatically one for the discrete algorithm. We develop contact Hamiltonian systems as a setting where the transfer can be made precise. A contact Hamiltonian H on J^1(\mathbbR^n) obeys the intrinsic decay identity \dot H = -H,\partial_s H , so an augmented energy \mathcalE built from H , together with the conformal rate \partial_s H , is a continuous-time rate certificate whenever \mathcalE controls the objective gap. Our main theorem states, under three named and independently checkable hypotheses, that an order- r contact splitting with step h transfers this certificate over the finite horizon set by backward error analysis. The discrete decay envelope is governed by the modified conformal factor up to O(h^r) perturbations plus a backward-error shadowing defect, and the mechanism is inherited exactly because the modified Hamiltonian is itself a contact Hamiltonian. Quadratic heavy ball is a fully solvable example: its projected dissipative-leapfrog spectrum agrees with established conformal-symplectic optimization theory, while the augmented contact Hamiltonian yields a sharp objective-to-certificate comparison that verifies the transfer hypotheses. For strongly convex objectives with state-dependent damping, an explicit Bregman-type Lyapunov certificate instead transfers by an auxiliary-shadowing corollary. The decomposition H=K+V+D into kinetic, objective-encoding potential, and dissipation terms serves as a design template, with a catalogue of closed-form sub-flows including contact-specific damping families. Numerical experiments confirm the predicted conformal-factor tracking orders and show competitive performance on ill-conditioned benchmarks and deep-learning tasks.

[LG-142] Distributed Convolutional Rank Regression over Decentralized Networks

链接: https://arxiv.org/abs/2607.23639
作者: Chunjing Li,Tiange Zhao,Xiaohui Yuan
类目: Methodology (stat.ME); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:This paper studies convolution rank regression (CRR) over decentralized distributed learning networks. We propose a novel decentralized CRR framework, in which estimators are obtained by solving consensus-constrained optimization with kernel-smoothed rank loss. The developed estimation scheme relies solely on local node data and information shared by neighboring nodes, thereby achieving privacy preservation and high communication efficiency. For heterogeneous network settings, we establish finite-sample error bounds for the decentralized CRR estimator and derive exact support recovery guarantees for the sparse decentralized CRR LASSO estimator. To facilitate numerical implementation, we adopt a generalized consensus ADMM to efficiently solve local subproblems across all network nodes. We verify the favorable performance of our developed approach via extensive numerical simulations and real-data experiments.

[LG-143] Learning switched non-linear dynamical systems from a single trajectory

链接: https://arxiv.org/abs/2607.23502
作者: Sunny G.W. Wang,Hemant Tyagi
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注: 56 pages, 2 figures

点击查看摘要

Abstract:We study empirical risk minimization for learning non-linear dynamical systems whose transition dynamics may switch over time. Under stability assumptions, and i.i.d switching over a set of K modes, we derive non-asymptotic bounds on the prediction risk expressed in terms of the metric entropy of the underlying function class. We instantiate our general result for Hölder and linear function classes, obtaining explicit convergence rates that depend on the effective sample size Tp_i , where T is the trajectory length and p_i is the probability of observing mode i . Numerical simulations support our theoretical findings. To the best of our knowledge, these results are the first non-asymptotic guarantees for learning switched nonlinear dynamical systems from a single trajectory.

[LG-144] wo-Timescale Hierarchical Reinforcement Learning for Resilient Operations

链接: https://arxiv.org/abs/2607.23434
作者: Young Hyun Cho,Franz Stoll,Will Wei Sun,Guang Lin,Stephan Biller
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注:

点击查看摘要

Abstract:Unexpected shocks recur in global operations, requiring decision rules that adapt as market and operating conditions change. Many operational systems also have hierarchical structures in which long-term and short-term decisions pursue a shared objective. We study how hierarchical reinforcement learning can strengthen resilience by adapting these interdependent rules jointly. We develop a two-timescale hierarchical reinforcement learning framework that adapts long-term and short-term policies at their respective time scales. Because the policies are interdependent, we synchronize their updates and prove, to our knowledge, the first convergence guarantees for coupled two-timescale learning. Over T periods, our policies’ average gap from an optimal policy pair is O(T^-1/2) , improving to O(\log T/T) when poor decisions produce clearer profit losses. In a used-car case study, inventory replenishment is the long-term decision and customer-arrival pricing the short-term decision. Relative to the strongest partially adaptive benchmark, the framework increases mean profit by 9.2% under joint demand-supply shocks and by 11.8% under a prolonged shock scenario, while maintaining a more stable profit trajectory over time. Short-term adaptation addresses routine seasonality and one-sided disruptions by responding immediately to changing conditions. Under joint demand-supply shocks, however, it is insufficient alone; long-term adaptation is also needed to create favorable conditions for short-term decisions. Joint adaptation thus yields higher and more stable profits through disruption and recovery. Because many organizations already use hierarchical planning, the framework strengthens operational resilience without altering existing decision structures.

[LG-145] Logit-Coordinate Generative Models for Mixed Continuous-Categorical Tabular Data

链接: https://arxiv.org/abs/2607.23348
作者: Yuefei Shen,Xiaotong Shen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Mixed continuous–categorical data pose a representation problem for continuous generative models. Flow Matching and Gaussian diffusion operate in Euclidean spaces, whereas categorical laws lie on probability simplices and may be highly imbalanced. We study a logit-coordinate framework that encodes categorical variables as smoothed natural parameters and combines them with transformed numerical variables. This yields common formulations of Logit Flow Matching and Logit Diffusion. We introduce a mixed-distribution discrepancy separating categorical marginal error from conditional continuous Wasserstein error, and derive stability bounds and imbalance-aware nonparametric rates linking vector-field or drift error to decoded mixed-distribution error. Controlled simulations show that scaled-logit coordinates improve or match one-hot coordinates, especially under severe rare-cell imbalance. Across four real-data benchmarks and ten splits per dataset, Logit FM improves the primary distributional metrics on three datasets and is comparable on Churn2; Block-Conditional Logit FM consistently improves the flat model; and Logit Diffusion generally improves over or matches One-Hot Diffusion.

[LG-146] Context-Adaptive Inference: A Unified Statistical and Foundation-Model View

链接: https://arxiv.org/abs/2607.23304
作者: Yue Yao,Caleb N. Ellington,Jingyun Jia,Baiheng Chen,Dong Liu,Rikhil Rao,Jiaqi Wang,Samuel Wales-McGrath,Yixin Yang,Zhiyuan Li,Eric P. Xing,Ben Lengerich
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 90 pages, 13 figures. Manuscript source and living version: this https URL

点击查看摘要

Abstract:Modern predictive systems are expected to adapt their behavior to the specific situation they are facing. A clinical model should not treat every patient the same; a retrieval-augmented model should change its answer when given different evidence; a mixture-of-experts model should route different inputs to different experts. We call this capability context-adaptive inference: before predicting, the system uses information about the current context to specialize its parameters or computation for that instance. This article provides a unified view of context-adaptive inference across three traditions that are usually treated separately: (i) explicit adaptation in statistics (e.g. varying-coefficient models, local regression, hierarchical sharing), (ii) rapid task-specific adaptation in meta-learning and transfer, and (iii) implicit adaptation in large foundation models via prompting, retrieval, and expert routing. We formalize these approaches under a common objective: to map context c to adapted parameters \theta© , then to predict via f(x; \theta©) . Under squared loss, linear prediction heads, and fixed features, we prove that explicit parameter adaptation and implicit routing are mathematically equivalent to kernel ridge regression on joint features of inputs and context. Building on this bridge, we propose practical design principles and evaluation metrics including adaptation-efficiency, routing stability, and context-specific robustness to guide when to specialize, how to constrain that specialization, and how to audit context-adaptive models in deployment. Finally, we identify open problems in identifiability, robustness under distribution shift, and efficient large-scale adaptation, outlining design principles for methods that are scalable, reliable, and transparent in real-world settings. Comments: 90 pages, 13 figures. Manuscript source and living version: this https URL Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME) Cite as: arXiv:2607.23304 [stat.ML] (or arXiv:2607.23304v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2607.23304 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-147] PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response Simulation

链接: https://arxiv.org/abs/2607.23293
作者: Shaoheng Xu,Chunyi Sun,Jihui Zhang,Amy Bastine,Prasanga N. Samarasinghe,Thushara D. Abhayapala
类目: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
*备注: Accepted for publication in the Proceedings of the 19th International Workshop on Acoustic Signal Enhancement (IWAENC 2026). Project page: this https URL

点击查看摘要

Abstract:Image-source-method (ISM)-based room impulse response (RIR) simulation is a useful and physically interpretable tool for acoustic scene modeling, but full-order ISM becomes computationally expensive as the reflection order and room complexity increase. We propose a physics-guided framework for fast RIR simulation that preserves the geometric structure of ISM while learning to retain only acoustically important image-source paths during online traversal. To recover energy removed by pruning, the proposed PathRIR uses a lightweight compensation multilayer perceptron to predict the missing late-tail energy envelope and generate a compensation tail whose energy follows that envelope. Experiments on irregular 3D rooms show that PathRIR reduces image-source computation and improves runtime efficiency over a full-order ISM simulator, while achieving low waveform- and decay-related errors. Ablation results show that adding the compensation tail improves waveform fidelity and reduces energy-decay-curve error, reverberation-time error, and direct-to-reverberant-ratio error, with modest runtime overhead.

[LG-148] Continuous surrogates versus threshold Boolean networks for modeling Arabidopsis ISR gene regulation

链接: https://arxiv.org/abs/2607.23289
作者: Gonzalo A. Ruz
类目: Molecular Networks (q-bio.MN); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注: To be published in IEEE CIBCB 2026

点击查看摘要

Abstract:Gene regulatory network modeling often requires balancing predictive accuracy and mechanistic interpretability. In this work, we compare continuous surrogate models and a discrete mechanistic model on the same \textitArabidopsis thaliana induced systemic resistance (ISR) dataset, using both the raw continuous gene-expression measurements and their sign-binarized representation. The study considers eight defense-related genes measured over nine time points and evaluates two continuous predictors, Random Forest (RF) regression and a Multi-Layer Perceptron (MLP), against a threshold Boolean network (TBN). The models are assessed using rolling-origin one-step prediction, recursive multi-step rollout, and interpretability analysis. RF achieved the best average one-step numerical performance in the continuous domain, with an MAE of 1.910 and an RMSE of 2.836, compared with 2.089 and 3.106 for the MLP. In the binary domain, the TBN obtained the best average one-step qualitative performance, with a binary accuracy of 0.550 and a Hamming distance of 3.600, compared with 0.500 and 4.000 for RF, and 0.495 and 4.040 for the MLP. In recursive rollout, the TBN exactly reproduced the observed binarized trajectory, while the MLP also showed near-perfect fidelity, with a trajectory binary accuracy of 0.986, and RF accumulated substantially larger deviation, with a trajectory binary accuracy of 0.708. These results highlight that local numerical accuracy and global qualitative dynamical fidelity are not necessarily aligned, and suggest that continuous surrogates and threshold Boolean networks should be viewed as complementary tools for modeling biological regulation.

[LG-149] Approximate reservoir computing with a semiconductor laser for reducing energy consumption

链接: https://arxiv.org/abs/2607.23288
作者: Tatsuki Ito,Kazutaka Kanno,Satoshi Kawakami,Atsushi Uchida
类目: Optics (physics.optics); Machine Learning (cs.LG); Chaotic Dynamics (nlin.CD)
*备注: 9 pages, 10 figures, 3 tables

点击查看摘要

Abstract:Photonic reservoir computing is a promising physical machine-learning technique for predicting time-series data. The quantization of the response signal from the reservoir is required for the implementation of photonic reservoir computing, and the number of quantization bits and sampling frequency need to be optimized to achieve high performance and low energy consumption. However, few studies have been reported to investigate the effect of bit quantization and sampling frequency. In this study, we introduce a concept of approximate reservoir computing with a semiconductor laser by quantizing the amplitude of node states in the reservoir and output weights. We evaluate the performance of a chaotic time-series prediction task and energy consumption per sample. We achieve significant reduction of energy consumption by optimizing the number of quantization bits, the sampling frequency, and the injection current of the semiconductor laser, while maintaining the prediction performance.

[LG-150] Learning Asymptotics with Convergence-Rate Guarantees using Linear Least Squares

链接: https://arxiv.org/abs/2607.23287
作者: Christos N. Efrem
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Combinatorics (math.CO); Numerical Analysis (math.NA)
*备注: 62 pages, 7 tables, 5 figures

点击查看摘要

Abstract:We introduce a new research area that is called Asymptotics Learning Theory (ALT) and combines optimization with asymptotic analysis. In particular, ALT provides a unified approach for computing unknown constants/parameters in proven asymptotic expansions using optimization theory. In this paper, we focus on a general asymptotic form which includes a broad class of asymptotics. Furthermore, we study two powerful numerical methods, namely, sliding Linear Least Squares (sLLSQ) and sliding Tikhonov Linear Least Squares (sT-LLSQ). For these techniques we rigorously prove asymptotic estimates that lead to sufficient conditions for convergence (to the correct values of unknown parameters) and convergence-rate guarantees. Despite their strengths, both methods have also limitations, e.g., slow convergence—or even, counterintuitively, divergence—in some cases. Moreover, we present fundamental applications in analytic combinatorics, a beautiful field of mathematics that deals with asymptotic enumeration of discrete structures using complex analysis. The proposed techniques complement existing approaches, such as the ratio method and its variants. Numerical examples also verify the theoretical results. Finally, we discuss interesting research directions in ALT.

[LG-151] Photonic reservoir computing with complex networks

链接: https://arxiv.org/abs/2607.23285
作者: Sion Park,Kohei Watabe,Satoshi Sunada,Tomoki Yamagami,Atsushi Uchida
类目: Optics (physics.optics); Machine Learning (cs.LG); Chaotic Dynamics (nlin.CD)
*备注: 18 pages, 10 figures, 3 tables

点击查看摘要

Abstract:Photonic reservoir computing has attracted increasing attention as a fast and low-cost approach for time-series prediction. Photonic reservoir computing utilizes the high speed, broad bandwidth, and spatial parallelism of light. However, the effect of the internal connection structure (network topology) on the computing performance has not been investigated for large-scale photonic reservoirs. In this study, we experimentally and numerically demonstrate photonic reservoir computing using a spatial light modulator to systematically evaluate the relationship between the network topology and the performance of reservoir computing. We introduce complex network structures such as small-world and scale-free network topologies of the internal nodes in the reservoir. We perform the memory capacity measurement and the one-step-ahead prediction task of the chaotic time series to compare the performance. We found that the small-world network exhibits the maximum memory capacity and the best prediction performance. Our numerical calculations reveal that the performance of the time-series prediction can be optimized by changing the rewiring probability of the network and the leak rate of the reservoir. We also implement photonic human brain network as a reservoir, which is designed by the connectomes of human brain activities. We found that the network topology strongly affects the performance of reservoir computing, and the small-world network structure outperforms the other configurations.

[LG-152] Beyond ICA: Identifiability by Symmetry Breaking

链接: https://arxiv.org/abs/2607.23182
作者: Pengzhou Wu
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:

点击查看摘要

Abstract:We prove the identifiability of deep generative models (DGMs) with piecewise-affine (PWA) decoders and Gaussian mixture model (GMM) priors, in a purely unsupervised setting. We introduce three algebraic contrast principles for symmetry breaking: domain contrast, which trivializes the mixture symmetry group; mechanism contrast, which ensures every decoder branch is witnessed by a unique boundary; and interaction contrast, which forbids parameter conspiracies between latent components and decoder branches. Together they exploit the interplay between the discrete combinatorics of the PWA map and the continuous symmetry structure of the latent GMM. Continuity is replaced by algebraic symmetry conditions; injectivity is decoupled from structural identification and required only for pointwise inversion. Our results form a hierarchy: from law identifiability (LID; latent distribution up to a global affine map) through map identifiability (MID; decoder up to the same map) to posterior and pointwise identifiability. The ICA-form ambiguity emerges under conditions on diagonal component covariances. Assumptions are only on the data-generating process, not on learning methods, except for the interaction contrast. To our knowledge this is the first to make algebraic symmetry-breaking the engine of nonlinear identifiability, the first to admit discontinuous decoders, and the first to handle fully non-injective decoders, where every observation admits multiple latent codes.

[LG-153] Adaptive Multi-Scale Forecasting and Gate-Localized Conformal Prediction for Multivariate Nonstationary Time Series

链接: https://arxiv.org/abs/2607.23165
作者: Ziling Ma,Junshu Jiang,Ángel López-Oriona,Ying Sun,Hernando Ombao
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME)
*备注:

点击查看摘要

Abstract:We propose ABF-T-GLCP, a model-agnostic framework for forecasting and uncertainty quantification in nonstationary multivariate time series. The central idea is to learn an adaptive predictive state representation for point forecasting and reuse it for conformal calibration. The forecasting module combines horizon-specific temporal experts through a learned gate and refines predictions using sparse predictive transfer across related series. The uncertainty module, Gate-Localized Conformal Prediction (GLCP), uses the learned gate state, together with temporal recency, to select locally relevant calibration residuals, thereby coupling uncertainty calibration to the predictive regimes used by the forecasting model. This shared representation allows point forecasts and prediction intervals to adapt consistently under evolving temporal dynamics while retaining the model-agnostic nature of conformal prediction and yielding approximate local coverage under mild stability conditions. Experiments on a large-scale high-frequency commodity forecasting benchmark show consistent gains in point forecasting accuracy and substantially narrower prediction intervals with empirical coverage close to the nominal level. Additional results indicate that the framework extends beyond the motivating financial application.

[LG-154] Operator Neural Jump ODEs: L2-optimal prediction in function spaces

链接: https://arxiv.org/abs/2607.23110
作者: Florian Krach,Oliver Löthgren,Josef Teichmann
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 25 pages

点击查看摘要

Abstract:In this paper, we study the extension of Neural Jump ODEs to infinite-dimensional function spaces. In particular, the underlying process X now takes values in L^2(\Xi, \mathbbR^d_X) instead of \mathbbR^d_X and the Operator NJ-ODE approximates the optimal predictor of this process by producing a representative of the conditional expectation. The NJ-ODE model is a framework for online learning the optimal prediction of continuous-time stochastic processes, given discrete, possibly irregular and incomplete past observations. In a series of works, this model has been extended to deal with generic path-dependent processes, with observation noise and dependent observations, with long-term predictions, and with input-output systems. However, throughout all of these works, the underlying processes were restricted to be finite-dimensional. In particular, function-valued problems, like yield curve or volatility surface predictions, could only be handled through discretization, which inherently leads to a loss of information. In this work, we build on ideas from Neural Operator methods that allow us to extend the NJ-ODE framework to an infinite-dimensional output process. To prove convergence of the NJ-ODE to the optimal prediction process, we develop a new approximation strategy that also generalizes previous works in the finite-dimensional setting by considerably weakening the assumptions.

[LG-155] Neural Network-Driven Volatility Drag Mitigation under Aggressive Leverag e

链接: https://arxiv.org/abs/2607.23068
作者: Christian Bongiorno,Efstratios Manolakis,Rosario Nunzio Mantegna
类目: Portfolio Management (q-fin.PM); Machine Learning (cs.LG)
*备注: 7 pages, 2 figures, 1 table. Published in ICAIF '25. Code: this https URL

点击查看摘要

Abstract:This paper introduces a compact reformulation of a modular end-to-end neural network for global minimum-variance portfolio optimization that decouples model complexity from both look-back window length and universe size. A five-parameter hyperbolic weighted moving average combined with a saturating exponential replaces the original 2,400-parameter lag-transformation layer, and a bidirectional gated-recurrent-unit eigencleaning module together with a streamlined marginal-volatility network reduce total learnable parameters from 39,586 to just 2,175. In out-of-sample tests against state-of-the-art nonlinear-shrinkage and risk-parity benchmarks, the compact network attains the lowest realized portfolio variance without compromising expected return. Under long-only constraints, the variance reduction supports substantially higher leverage while maintaining comparable drawdown control. Validation in a high-fidelity trading simulator that incorporates realistic margin-call dynamics confirms enhanced over-leverage resilience. These findings demonstrate that end-to-end variance-minimization architectures can achieve substantial parameter efficiency and robust capital-efficiency gains without sacrificing risk-adjusted performance.

[LG-156] Characterizing Arbitrary Lindbladian Dynamics with a Few Pauli Measurements

链接: https://arxiv.org/abs/2607.23044
作者: Taiqi Zhou,Weiyuan Gong
类目: Quantum Physics (quant-ph); Information Theory (cs.IT); Machine Learning (cs.LG)
*备注: 24 pages, 1 figure

点击查看摘要

Abstract:Quantum devices are open systems whose dynamics interleave coherent evolution with dissipation, and benchmarking, error mitigation, and error correction all rest on a faithful model of both. Existing characterization protocols either assume prior knowledge of the interaction and noise structure, or demand ancillas, entangled probes, or mid-circuit control, or capture only the Pauli-diagonal part of the noise. Here, we present a protocol that reconstructs an arbitrary sparse Markovian generator, including every Hamiltonian together with the jump operator coefficients, using only product Pauli state preparation, single uninterrupted forward evolutions, and product Pauli measurements. Given a sparsity budget M_0 and a strength bound \Gamma of the Lindbladian, every coefficient is learned to precision \epsilon from \widetildeO(\Gamma^2M_0^2/\epsilon^4) experiments and \widetildeO(\Gamma M_0^2/\epsilon^2) total evolution time, with both supports identified from data without locality assumptions. The protocol runs at a logarithmic number of positive evolution times on a hardware clock lattice and is provably robust to calibrated state-preparation and measurement errors.

[LG-157] Covariance-Boosted Gaussian Processes for Spatiotemporal Irregularities

链接: https://arxiv.org/abs/2607.23018
作者: Jeremy Ovadia
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Space Physics (physics.space-ph); Methodology (stat.ME)
*备注:

点击查看摘要

Abstract:Nonstationary Gaussian process (GP) models are powerful tools for capturing input-dependent variability by adapting to observed data. However, with limited sampling and highly parameterized covariance structure, they are often prone to overfitting and overconfident uncertainty estimates, potentially leading to misleading predictions in safety-critical applications. Motivated by ionospheric modeling for satellite-based augmentation systems (SBAS), this paper proposes a Covariance-Boosted Gaussian Process (CBGP) framework centered upon boosting covariance priors to discover nonstationary latent functions for signal and observation variation that capture irregularities in the input domain. An additional layer of GP modeling of “partially-whitened” observations guides latent function relative error estimation that is used to iteratively update weak priors in a gradient descent-like procedure. Following boosting, restrictions are imposed upon prior covariances to prevent overfitting while posterior uncertainties are inflated to prevent model overconfidence. CBGP model efficacy and robustness are demonstrated through out-of-sample testing of both simulated and real-world applications that meet a three-nines integrity standard. The modeling of an extensive ionospheric storm dataset over South America suggests accurate and reliable means to compute SBAS ionospheric corrections in the most challenging space weather environment using regional models that are more informed and responsive than local fitting performed by currently-operating SBAS.

[LG-158] Nesterov acceleration in optimizing over probability measures

链接: https://arxiv.org/abs/2607.23008
作者: Jiaqi Tang,Qin Li,Wilfrid Gangbo
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Optimization over probability measures has become an increasingly important paradigm in modern machine learning, scientific computing, and uncertainty quantification. Motivated by Nesterov’s accelerated gradient method in Euclidean space, we develop Heavy-ball and Nesterov acceleration methods over the probability measure space \mathcalP_2 and establish non-asymptotic convergence guarantees that match their Euclidean counterparts. In particular, we derive convergence rates with respect to both the number of iterations and the number of particles used to represent the underlying probability distributions. Extending accelerated optimization from Euclidean space to probability measures is challenging. The natural notion of momentum requires concepts such as tangent bundles of the set of probability space and they are hard to operate numerically. To overcome these difficulties, we introduce two complementary lifting procedures. The first lifts probability measures to phase space through a Hamiltonian formulation, introducing momentum variables into the dynamics. The second lifts probability measures to a common Hilbert space, restoring the linear structure required for convergence analysis while simultaneously yielding executable particle dynamics. Together, these two complementary lifting procedures provide a systematic methodology for designing, analyzing, and implementing momentum-based accelerated optimization methods over probability measure spaces. Subjects: Optimization and Control (math.OC); Machine Learning (cs.LG) Cite as: arXiv:2607.23008 [math.OC] (or arXiv:2607.23008v1 [math.OC] for this version) https://doi.org/10.48550/arXiv.2607.23008 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-159] Robust Conformalized Selection with Noisy Responses

链接: https://arxiv.org/abs/2607.22985
作者: Chengyao Yu,Hongxin Wei,Bingyi Jing
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注:

点击查看摘要

Abstract:Conformalized selection has been widely applied to select high-quality candidates from large datasets with rigorous uncertainty quantification, such as reliable labeling, drug discovery, and the alignment of large language models. Nevertheless, existing methods assume clean responses on calibration data, an assumption that rarely holds in practice. In this paper, we formulate the above tasks as selecting candidates with true predicted labels or with responses exceeding certain values. We demonstrate that existing conformal selection methods fail to control the false discovery rate (FDR) or suffer from severe power loss under contaminated calibration data. To that end, we propose Robust Conformalized Selection (RCS), a unified framework for selective classification with valid FDR control under general label contamination. The key insight of RCS lies in a novel statistical reduction: by separately conditioning on different classes, we translate the intractable label noise into a localized covariate shift problem, which then enables a covariate-adjusted empirical-Bayes-type estimate of the number of false selections. Statistical properties such as the asymptotic FDR control, power optimality, and robustness of RCS are established. We further develop an instantiation of RCS under randomized response model, and also apply RCS to the task of selecting candidates with large response values. Extensive experiments on both simulated and real-world datasets demonstrate the effectiveness of RCS.

[LG-160] Variable Importance Identification Through Lazy Training for Binary Classification

链接: https://arxiv.org/abs/2607.22979
作者: Anand Singh,Luke Pennella,Eshan Kabir,Xiaoxi Shen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Deep neural networks have been widely used in many applications (e.g., computer vision and natural language processing); however, understanding their explainability remains a challenging task. Recently, substantial research has been devoted to improving the explainability of deep neural networks, with most of this work focusing on the regression framework. In this paper, we instead focus on the binary classification framework and adopt a variable-importance framework combined with the idea of lazy training to propose an efficient algorithm for identifying important features. From a theoretical perspective, our method relies on only a minimal set of assumptions and achieves well-controlled error rates. The validity of the proposed method and algorithm is examined through extensive simulation studies and real-data applications.

[LG-161] On the Order-Conditional Optimality of Gaffkes Bound

链接: https://arxiv.org/abs/2607.22971
作者: George Bissias,Erik Learned-Miller
类目: atistics Theory (math.ST); Machine Learning (cs.LG); Probability (math.PR); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Let X = (X_1, \ldots, X_n) be a random vector from any Borel probability law on \mathbbR_+^n . We revisit the problem of deriving a lower confidence bound (LCB) on a scalar parameter of that law. We recast classical work, beginning with Buehler, in purely probabilistic terms to form a more accessible and extensible framework. We then specialize the framework to the case where the components of X are independent. In this context, we prove that Gaffke’s bound is Buehler optimal for the order that it induces with respect to the maximum marginal mean parameter: max_i \in [n] E_Q[X_i] , which reduces to the common mean when the X_i are independent and identically distributed. That is to say, no other valid LCB that orders samples in the same way as Gaffke’s bound can improve on it with respect to this parameter.

[LG-162] Amortized Bayesian Causal Discovery of Extended Factor Graphs

链接: https://arxiv.org/abs/2607.22934
作者: Yichen Gu,Yuxuan Song,Weizhou Qian,Yixin Wang,Joshua Welch
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Molecular Networks (q-bio.MN)
*备注:

点击查看摘要

Abstract:Learning causal graphs from interventional data is a challenging problem with broad applications. In molecular biology, for example, a central goal is to uncover gene regulatory networks from large-scale perturbation data. An ideal algorithm for this task should scale to thousands of nodes, incorporate interventions even when their targets are unknown, quantify uncertainty, and provide identifiability guarantees. However, existing approaches—e.g. approaches using score-based optimization or approximate Bayesian inference—often fail to meet all of these criteria. To address these limitations, we develop Amortized Bayesian Causal Discovery of Extended Factor Graphs (ABCDEFG). Our method guarantees exact acyclicity, scales to graphs with thousands of nodes, and naturally handles interventions even when their targets are unknown. Additionally, ABCDEFG estimates a posterior distribution whose maximum a posteriori estimate provably identifies the true causal graph up to an equivalence class. On simulated datasets, ABCDEFG achieves state-of-the-art accuracy, producing a well-calibrated posterior distribution while outperforming previous score-based and approximate Bayesian methods. Applied to large-scale single-cell perturbation data, ABCDEFG identifies both established and novel gene targets of growth factors.

[LG-163] Practical advantage beyond the quadratic speedup limit with fully-quantum walks

链接: https://arxiv.org/abs/2607.22818
作者: Massimiliano Incudini,Guglielmo Mazzola
类目: Quantum Physics (quant-ph); Disordered Systems and Neural Networks (cond-mat.dis-nn); Machine Learning (cs.LG)
*备注: Data and notebooks to reproduce the results are publicly available

点击查看摘要

Abstract:We introduce a new class of fully-quantum Metropolis walks in which both the proposal and acceptance steps are intrinsically quantum. Unlike standard quantum walks obtained by quantizing classically efficient Markov chains, our algorithm employs Hamiltonian simulation as a quantum-native proposal mechanism, enlarging the class of quantum walks beyond classical counterparts. We target the problem of sampling from the low-temperature Gibbs distribution of classical dense Ising models, within a fixed error in total variation distance. This approach achieves about a cubic polynomial asymptotic advantage over previous quantum-walks, resulting in a total sixth-degree polynomial queries speedup compared to the best classical walk. This shows that speedups beyond the widely assumed quadratic limit are possible within the quantum walk formalism. We perform a complete fault-tolerant compilation of all algorithmic primitives and benchmark against CPU, GPU, and FPGA implementations of the best classical Markov chain. Under identical hardware assumptions, the resulting advantage runtime crossover is reduced from approximately 10^3 years for conventional quantum walks to less than one day. These results identify fully-quantum Markov chains as a promising route toward practical quantum advantage.

[LG-164] Subject-Level Heterogeneity in EEG Motor Imagery Decoding: A Large-Scale Benchmark and Portfolio-Based Reduction of the Search Space

链接: https://arxiv.org/abs/2607.22778
作者: Paul Barbaste,Olivier Oullier,Xavier Vasques
类目: Neurons and Cognition (q-bio.NC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:Robust EEG motor imagery decoding remains limited by strong inter-individual variability, making it difficult to identify pipelines that generalize across users. We present a large-scale, standardized within-session benchmark of decoding pipelines across three public datasets: Cho2017 (52 subjects), PhysionetMI (109 subjects), and Zhou2016 (4 subjects). Using a common MOABB LeftRightImagery setting, two frequency bands (8-15 Hz and 8-30 Hz), and a broad combination of feature extraction, preprocessing, and classification steps, we analyzed 216,714 raw evaluation rows, which after structured aggregation yielded 44,928, 109,000, and 4,192 subject-level observations respectively. Covariance tangent-space projection (cov-tgsp) and Common Spatial Patterns (CSP) consistently defined the strongest methodological families, though their relative ordering was dataset-dependent. On Cho2017, the best family-level mean accuracy came from cov-tgsp in 8-30 Hz (0.712 +/- 0.140), whereas Zhou2016 favored CSP (0.832 +/- 0.121 in 8-15 Hz). These aggregate rankings concealed substantial subject-level heterogeneity: 42 distinct winning pipelines across 52 Cho2017 subjects, and 93 across 109 PhysionetMI subjects. We then used the benchmark as an empirical performance landscape for building compact portfolios of pipelines of size K. Several construction procedures were compared, including a ranking-based Top-K Mean heuristic and search-based strategies. Results were broadly consistent, with Top-K Mean giving the best trade-off. A single best global pipeline already retained 94.2% of the oracle in Cho2017 and 81.8% in PhysionetMI; at K = 12, oracle retention rose to 96.5% and 90.0%. The landscape is therefore subject-dependent, and this heterogeneity can be exploited through compact portfolios that make personalization more feasible. Subjects: Neurons and Cognition (q-bio.NC); Machine Learning (cs.LG) Cite as: arXiv:2607.22778 [q-bio.NC] (or arXiv:2607.22778v1 [q-bio.NC] for this version) https://doi.org/10.48550/arXiv.2607.22778 Focus to learn more arXiv-issued DOI via DataCite (pending registration)

[LG-165] LRNet: Estimating Individual Treatment Effect based on Local Information and Single Learner Structure

链接: https://arxiv.org/abs/2607.22762
作者: Ali Haghpanah Jahromi,Mohammad Taheri,Zohreh Azimifar
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 7 pages, 1 figure, 2 tables. This is the author-accepted manuscript of the paper presented at the 2nd International Conference on Artificial Intelligence and Software Engineering (AI-SOFT 2024), Shiraz University, Iran

点击查看摘要

Abstract:Causal inference has become a central issue across various fields, including computer science, statistics, economics, education, healthcare, and medicine. The broad applicability of this discipline has garnered increased research funding and attention. In recent years, the estimation of causal effects from observational data has gained traction due to the vast amounts of collected data and the lower costs compared to randomized controlled trials. Advances in causal effect estimation methods have enhanced service personalization tools. For instance, these tools can help identify the most effective type of treatment (considering both cost and success rate) for each patient among different medical service options. This paper proposes an innovative method for estimating the heterogeneity of treatment effects. The structure of the proposed model is based on a deep neural network and a pseudo-single learner. The proposed method has been compared with other state-of-the-art methods on the IHDP benchmark. Acceptable results have been obtained by using one estimator to estimate the potential outcomes of two treatment groups. Accordingly, this paves the way for further development and improvement of the proposed method.

[LG-166] A Resolution of the SS–RS–GD Inequalities

链接: https://arxiv.org/abs/2607.22620
作者: Binghui Peng
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:

点击查看摘要

Abstract:Yun, Sra, and Jadbabaie (COLT 2021, open question) conjectured the SS–RS–GD inequalities: for well-conditioned symmetric matrices A_1,\dots,A_n , the operators W_ss , W_rs , and W_gd that encode the expected iterate of single-shuffle SGD, random-reshuffle SGD, and gradient descent on a quadratic finite sum should satisfy [ |W_ss|\le | W_rs|\le |W_gd|. ] The conjecture is resolved, \bullet SS-RS inequality fails. Already for n=3 , K=2 , and d=4 , we exhibit explicit PSD matrices whose condition number is arbitrarily close to 1 , yet |W_ss||W_rs| . \bullet RS-GD inequality holds. For every symmetric A_i with \bigl(1-\frac14n^2+1\bigr)I\preceq A_i\preceq I , one has |W_rs|\le|W_gd| . The proof was found via GPT-5.5 Pro extended prompted by the author. Subjects: Optimization and Control (math.OC); Machine Learning (cs.LG); Machine Learning (stat.ML) Cite as: arXiv:2607.22620 [math.OC] (or arXiv:2607.22620v1 [math.OC] for this version) https://doi.org/10.48550/arXiv.2607.22620 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Binghui Peng [view email] [v1] Tue, 16 Jun 2026 16:17:42 UTC (10 KB) Full-text links: Access Paper: View a PDF of the paper titled A Resolution of the SS–RS–GD Inequalities, by Binghui PengView PDFHTML (experimental)TeX Source view license Current browse context: math.OC prev | next new | recent | 2026-07 Change to browse by: cs cs.LG math stat stat.ML References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from

[LG-167] DRP-FLR: Data-Driven Assessment of Demand Response Potential for Flexible Load Regulation in Smart Grids

链接: https://arxiv.org/abs/2607.22590
作者: Yunhao Yao,Siyu Jing,Yang Yang,Qiang Xu,Changqi Weng,Xiang-Yang Li
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注: 10 pages, 6 pages

点击查看摘要

Abstract:The rapid growth of AI workloads and renewable energy resources exacerbates supply-demand imbalance in power systems, making traditional load regulation designed for efficient allocation inadequate and motivating demand response (DR) mechanisms to enable load controllability in smart grids. However, existing DR-oriented approaches either focus on optimizing electricity cost or occupant comfort with limited benefit to system-level balance. Others overlook the diverse and dynamic consumption patterns of heterogeneous energy entities, leading to significant over- or under-regulation. Therefore, we propose DRP-FLR. First, DRP-FLR achieves accurate short-term load forecasting by embedding exogenous knowledge (e.g., entity information, prediction time) into historical load representations. Next, it constructs entity-specific load-pattern profiles by clustering historical load curves, and estimates DR potential by matching forecasted loads with pattern profiles. Finally, DRP-FLR formulates flexible load regulation as a mixed-integer optimization problem and solves it with an MILP solver to jointly optimize DR utilization, participant economic benefit, and renewable accommodation, while enforcing supply-demand balance and economic feasibility. Experiments on a regional grid and a campus microgrid show that DRP-FLR reduces regulation deviation by 36.63%-91.87% and improves participant benefit by 44.66% on average.

[LG-168] Learning to Optimize at Scale: A Benders Decomposition-TransfORmers Framework for Stochastic Combinatorial Optimization

链接: https://arxiv.org/abs/2607.22550
作者: Seung Jin Choi,Kimiya Jozani,Josh Cooper,Esra Buyuktahtakin Toy
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:

点击查看摘要

Abstract:We propose a learning-augmented Benders decomposition framework to solve large-scale two-stage stochastic mixed-integer programs. We focus on the two-stage stochastic capacitated lot-sizing problem (TSSCLSP) under demand uncertainty. Our method accelerates the convergence of the decomposition by using a pre-trained TransfORmer model to rapidly generate high-quality approximate solutions for the scenario subproblems. This hybrid strategy uses the TransfORmer predictions to generate strong optimality and feasibility cuts, effectively guiding the Benders master problem. Our framework includes a novel expandable generation mechanism, allowing a model trained on a fixed horizon to solve instances of arbitrary length. For the test set considered, our method solves instances up to T = 270, a scale previously intractable for this approach, while maintaining zero infeasibility in the generated subproblem solutions. This demonstrates the potential of TransfORmers as powerful surrogate solvers embedded within classical decomposition algorithms.

附件下载

点击下载今日全部论文列表