Personalized Daily Arxiv Papers 09/04/2026

This project is adapted from tatsu-lab/gpt_paper_assistant. The source code of this project is at Variante/gpt_paper_assistant

About me on Bilibili. Help keep the website running:

Topics

Paper selection prompt and criteria (jump to the section by clicking the link):

1. Application of diffusion models and vision-language models (VLMs) to robot manipulation.

2. New methodological improvements to self-supervised learning (SSL) for image or video representation.

3. Video segmentation using unsupervised or self-supervised methods.

4. Transfer learning across modalities (e.g., audio-to-video, optical flow, language-to-video) for improved video understanding.

5. Advances in 3D generation using generative models, including image-to-3D and text-to-3D.

6. Recent progress in 3D reconstruction and generation with Gaussian Splatting, NeRF, or mesh generation.

Go beyond


Topic 1

1000. Scaling Bimanual Household Manipulation from 1,500 hours of Demonstrations to On-Policy Corrections [more]
Authors: Jiafeng Xu, Qi Li, Yan Shen, Yiyu Ren, Travis Davies, Shaowen He, Ze Wang, Yifan Yang, Ran Cheng, Hao Dong

1001. WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models [more]
Authors: Chenhao Zhang, Hanyu Zhao, Hang Cheng, Tengfei Pan, Long Zeng

1005. Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies [more]
Authors: Yue Yang, Diego Romeres, Chiori Hori, Gedas Bertasius, Daniel Szafir, Siddarth Jain

1006. Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis [more]
Authors: Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang

1008. Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations [more]
Authors: Onat \c{S}ahin, Mohammad Altillawi, George Eskandar, Carlos Carbone, Ziyuan Liu

1012. Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning [more]
Authors: Muyuan Liu (GENISOM AI, Beijing, China), Yue Huang (GENISOM AI, Beijing, China), Zheng Liang (GENISOM AI, Beijing, China), Xiang Gao (GENISOM AI, Beijing, China)

1013. FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation [more]
Authors: Yutian Zhang, Siyuan Ma, Liwen Yang, Yang Li, Ce Hao, Haozhen Chi, Dong We, Qiaojun Yu, Dibo Hou

1019. Latent Energy Action Planning with World Models [more]
Authors: Phu Pham, Aniket Bera

Back to [top]


Topic 2

2010. The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs [more]
Authors: Yumeng Shi, Quanyu Long, Yin Wu, Wenya Wang

2018. Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision [more]
Authors: Shravan Venkatraman, Wenshuai Zhao, Mohammad Hassan Vali, Arno Solin

Back to [top]


Topic 3

Back to [top]


Topic 4

4021. Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning [more]
Authors: Ye-Chan Kim, Seunghee Choi, SeungJu Cha, Si-Woo Kim, Hwiseon Kim, Hyungee Kim, Dong-Jin Kim

Back to [top]


Topic 5

5002. Sparse auto-regressive modeling for scene generation from multi-view images [more]
Authors: Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel, Wonjune Cho, Bardienus Pieter Duisterhof, Vincent Leroy, Jerome Revaud

5003. VisCAD: A Foundation Model Suite with Multimodal Industrial CAD Intelligence [more]
Authors: JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong, Zhichao Huang, Guanlin Li, Hongsen Liu, Ziqi Liu, Yichen Long, Luya Wang, Yuchen Wang, Wenxiang Wu, Huimu Yu, Ning Zhang

5009. RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents [more]
Authors: JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong, Zhichao Huang, Guanlin Li, Zongzhen Li, Hongsen Liu, Yichen Long, Wei Wang, Yuchen Wang, Dongyue Yang, Huimu Yu, Xianwen Zhong

5011. Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States [more]
Authors: Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy

5017. Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations [more]
Authors: Denis M. Akola, David F. Fouhey

5020. OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping [more]
Authors: Zelong Lv, Sicheng Xu, Jianfeng Xiang, Ruicheng Wang, Yue Dong, Yu Deng, Guangzhong Sun, Jiaolong Yang

Back to [top]


Topic 6

6004. Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction [more]
Authors: Chin-Yang Lin, Yang-Che Sun, Cheng Sun, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Wei-Chen Chiu, Yu-Lun Liu

6007. STARS-GS: Structure-Aware Regularized Gaussian Splatting for Large-Scale Aerial Surface Reconstruction [more]
Authors: Bocheng Li, Wenjuan Zhang, Jie Pan. Dongxu Han, Xuesong Ma, Yiling Yao, Yaning Wang

6014. Laplacian Frequency Hierarchies for Efficient 3D Gaussian Splatting Training [more]
Authors: Yixiong Yang, Sisheng Zhang, Qingsong Yan, Shaohuai Shi, Qiang Wang

6015. TruncGradGS: Improved 3D Gaussian Splatting via Truncated Gradient Updates [more]
Authors: Theo Morales, Nhat-Quynh Le-Pham, Robin Atkins, Binh-Son Hua

6016. Stable and Scalable Bundle Adjustment of Holistic 3D Structures [more]
Authors: Shaohui Liu, R'emi Pautrat, Daniel Barath, Richard Hartley, Viktor Larsson, Marc Pollefeys

6022. VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues [more]
Authors: Ernesto Lozano, Alberto Jaenal, Javier Civera

Back to [top]


Go beyond

Back to [top]


Full paper list

1000. Scaling Bimanual Household Manipulation from 1,500 hours of Demonstrations to On-Policy Corrections

ArXiv: 2609.03591 [page] [pdf]

Authors: Jiafeng Xu, Qi Li, Yan Shen, Yiyu Ren, Travis Davies, Shaowen He, Ze Wang, Yifan Yang, Ran Cheng, Hao Dong

Abstract: Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality large scale human demonstration data. In this work, we release 1,500 hours of diverse bimanual manipulation demonstrations covering everyday household tasks, and use this comprehensive corpus to train XR-2, a powerful vision-language-action (VLA) model. Enabled by a purpose built high throughput data pipeline and a carefully designed multi stage training paradigm, XR-2 attains strong manipulation performance in our systematic experiments while retaining favorable training efficiency and high data utilization. We further study two critical scaling axes: varying the amount of expert demonstration data, and post training on DAgger correction data from real time human interventions. In both settings, task success rate improves steadily over the data ranges we probe, exhibiting a clear consistent scaling trend at our current data scale. These results validate both the learning capacity of XR-2 and the promising scaling properties of the released dataset, which we open source to support reproducible research on bimanual robot manipulation learning.

Comment: Criterion 1: trains XR-2, a vision-language-action model for bimanual household manipulation, from 1,500 hours of demonstrations plus DAgger-style on-policy corrections with clear scaling trends in success rate.

Relevance: 10 Back to [topic] [top]

1001. WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

ArXiv: 2609.03681 [page] [pdf]

Authors: Chenhao Zhang, Hanyu Zhao, Hang Cheng, Tengfei Pan, Long Zeng

Abstract: Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinforcement learning with expensive and potentially unstable real-world exploration. World models offer a promising alternative by evaluating candidate behaviors through imagined futures, yet effective post-training requires more than accurate prediction: imagination must be scheduled where it is useful, bounded within reliable horizons, and translated into trustworthy policy supervision. In robotic manipulation, the value of imagination varies substantially across execution stages, while extended rollouts can accumulate prediction errors and introduce unreliable learning signals. We introduce WISE (World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models), a unified framework that coordinates when and how world-model imagination is used during policy refinement. WISE selectively invokes imagination at interaction-relevant states, performs bounded multi-view rollouts, evaluates candidate futures using progress and completion signals, and uses their relative outcomes to refine actions generated from real interaction contexts. Extensive experiments with both $\pi_0$ and $\pi_{0.5}$ demonstrate consistent improvements across diverse manipulation tasks while reducing GPU computation time by approximately 80% compared with full imagination. Real-world evaluations further show substantial gains in robustness and generalization under diverse real-world distribution shifts.

Comment: Criterion 1: WISE uses world-model-guided imagination scheduling to post-train VLA manipulation policies, improving both π0 and π0.5 while reducing GPU computation by about 80% versus full imagination.

Relevance: 10 Back to [topic] [top]

1005. Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies

ArXiv: 2609.03142 [page] [pdf]

Authors: Yue Yang, Diego Romeres, Chiori Hori, Gedas Bertasius, Daniel Szafir, Siddarth Jain

Abstract: Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous robot demonstrations encourages spurious inter-sensor correlations rather than task-relevant signal, a failure we term modality entanglement. Under real-world occlusions and distractors, this manifests as nuisance sensitivity to corruption of uninformative sensors and single-modality insufficiency when only one informative sensor remains intact. We propose Evidence-Gated Regularization (EGR), a modality-agnostic training objective that introduces zero inference-time overhead. EGR derives a per-frame and per-sensor task-relevance signal to gate two state-conditional consistency objectives: invariance on low-evidence sensors, and single-sensor sufficiency on high-evidence ones. We introduce a benchmark based on BEHAVIOR-1K, comprising a fast inference-only diagnostic suite and 47 rollout-based skills targeting modality entanglement. We validate EGR on this benchmark and on two real-robot setups with fundamentally different embodiments: a bi-manual setup with two Kinova arms and three RGB cameras, and a single-arm MELFA ASSISTA setup combining vision and GelSight tactile sensors. EGR improves simulation success rates (SR) from 12.5% to 16.4% under full modalities (+31%), from 9.4% to 16.5% under uninformative-sensor corruption (+75%), and from 2.8% to 6.1% under single-sensor fallback (+120%). Under physical-object distractors, EGR boosts SR from 30% to 85% on the bi-manual setup (+183%) and from 55% to 70% on the tactile setup (+27%).

Comment: Criterion 1: proposes Evidence-Gated Regularization for robust VLA robot manipulation policies, evaluated on a BEHAVIOR-1K rollout benchmark and two real-robot setups with gains such as 30% to 85% success under bimanual distractors.

Relevance: 9 Back to [topic] [top]

1006. Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

ArXiv: 2609.04096 [page] [pdf]

Authors: Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang

Abstract: This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/

Comment: Criterion 1: AdaRoboVLG is a vision-language-grasp framework that integrates composable foundation-model priors into grasp synthesis and demonstrates cross-hand generalization in simulation and real-world grasping.

Relevance: 9 Back to [topic] [top]

1008. Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations

ArXiv: 2609.03657 [page] [pdf]

Authors: Onat \c{S}ahin, Mohammad Altillawi, George Eskandar, Carlos Carbone, Ziyuan Liu

Abstract: 3D scene representations like NeRF and 3D Gaussian Splatting (3DGS) suffer severe artifacts in sparse-view settings. Recent generative 3D artifact fixers attempt to address this, but rely on paired corrupted and clean renders requiring costly, per-scene reconstructions across varying view configurations. While 2D image augmentations act as instant regularizers, no explicit equivalents exist for 3D representations to preserve spatial consistency across views, an essential property for 3D-aware training. We propose 3D Morphological Perturbations as an optimization-free regularizer that preserves spatial consistency. Leveraging explicit 3DGS, we treat each Gaussian as a fundamental building block - analogous to a 2D pixel - and apply perturbations across its morphological parameter space via scale, rotation, and pruning. Our method eliminates per-scene 3DGS optimization loops from dataset curation while enabling models to learn stronger geometric priors than sparse-view baselines in diagnostic ablations conducted on a lightweight video diffusion sandbox. Scaled to a 14B-parameter video model via ControlNet, our approach maintains visual fidelity while reducing mean depth error by 12.5% over state-of-the-art image-to-image 3D artifact refiners, ultimately boosting downstream robotics policy success rates by up to 8.0% across 3 of 4 manipulation tasks.

Comment: Criterion 1: uses 3D Gaussian morphological perturbations to train a video diffusion/ControlNet 3D artifact refiner, reducing mean depth error by 12.5% and improving downstream robot manipulation policy success by up to 8.0%.

Relevance: 8 Back to [topic] [top]

1012. Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

ArXiv: 2609.03565 [page] [pdf]

Authors: Muyuan Liu (GENISOM AI, Beijing, China), Yue Huang (GENISOM AI, Beijing, China), Zheng Liang (GENISOM AI, Beijing, China), Xiang Gao (GENISOM AI, Beijing, China)

Abstract: Action-conditioned JEPA world models enable planning toward visually specified goals without reconstructing future pixels, yet latent prediction alone does not explicitly encourage the learned representations to retain information relevant to robotic control. We introduce an end-to-end JEPA world model that augments latent prediction with inverse dynamics (IDM) and state alignment (SA). While inverse dynamics discourages latent collapse and makes latent transitions informative of the actions that produced them, state alignment grounds consecutive representations in their associated physical configuration and motion. Across four benchmark tasks, our model attains the highest success rates on TwoRoom (100%), PushT (98%), and OGBench-Cube (87%), while performing comparably to LeWorldModel on Reacher. Our ablation further shows that adding state alignment consistently improves planning success over IDM alone across all four tasks. Although LeWorldModel, our primary baseline, attains higher average straightening on OGBench-Cube, transition-subspace analysis shows that its transition energy is concentrated in a substantially lower-dimensional subspace. Our state-aligned model exhibits a higher effective transition dimension than LeWorldModel and improves planning over IDM alone, supporting state alignment as an effective complement to inverse dynamics for robotic planning.

Comment: Criterion 1: introduces an action-conditioned JEPA world model with inverse dynamics and state alignment for goal-conditioned robotic planning, achieving results such as 98% success on PushT and 87% on OGBench-Cube.

Relevance: 8 Back to [topic] [top]

1013. FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation

ArXiv: 2609.03889 [page] [pdf]

Authors: Yutian Zhang, Siyuan Ma, Liwen Yang, Yang Li, Ce Hao, Haozhen Chi, Dong We, Qiaojun Yu, Dibo Hou

Abstract: Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction control. Existing Vision-language-action (VLA) models generate task-level actions from visual and linguistic observations, but cannot interpret the physical interactions induced by those actions. While the whole-body control (WBC) policy can stabilize the robot, it cannot distinguish task-relevant interaction forces from forces induced by external disturbances during manipulation. Although force/torque sensors provide direct measurements of physical interactions, retrofitting them entails additional hardware costs and substantial integration effort, particularly for platforms not designed with sensor integration in mind. To address this problem, we propose FWBC-VLA, a force-aware framework that bridges task-level VLA action generation and low-level whole-body compensation control for wheeled-legged robots. First, we introduce HSR-Force, a sensorless residual-torque estimator for inferring contact strength and its temporal variation. These contact estimates are then encoded as tokens and injected into the VLA action expert during action decoding, enabling the policy to perceive contact onset, sustained loading, and release. For loco-manipulation tasks, all parameters of the pretrained VLA backbone are fine-tuned on our WL\&Arm Dataset, which comprises more than 5,000 episodes. Moreover, the robot’s proprioceptive state, the Jacobian-derived body-frame force estimate, and the estimated contact state are jointly fed into a compensation generator to produce corrective actions. The manipulation-centric actions are subsequently combined with the corrective actions and passed to the WBC policy for execution. Real-world experiments on whiteboard wiping and door opening with a door closer demonstrate the effectiveness of our FWBC-VLA in contact-rich loco-manipulation.

Comment: Criterion 1: FWBC-VLA fine-tunes a pretrained VLA backbone with force/contact tokens and whole-body compensation for contact-rich manipulation tasks such as whiteboard wiping and door opening.

Relevance: 8 Back to [topic] [top]

1019. Latent Energy Action Planning with World Models

ArXiv: 2609.03294 [page] [pdf]

Authors: Phu Pham, Aniket Bera

Abstract: Latent world models support efficient model predictive control from high-dimensional observations, yet optimizing a single learned latent objective can favor action sequences whose decoder-predicted terminal descriptor does not match the goal descriptor. We introduce Latent Energy Action Planning (LEAP), which treats the complete action horizon as a differentiable variable and optimizes it through a frozen LeWorldModel (LeWM). LEAP couples terminal latent goal matching with a terminal-window state energy. Low energy requires the predicted terminal latent to agree with the goal latent and the decoder-predicted terminal descriptor to agree with the goal descriptor. A frozen goal-conditioned proposal initializes the search, a quasi-Newton solver refines actions through the autoregressive rollout, and post-optimization projection enforces the admissible action range. Across four control domains using the officially released LeWM checkpoints, the complete LEAP planning system raises mean success from 77.5% for LeWM planned with the cross-entropy method (LeWM+CEM) to 94.8% under a matched protocol, a 17.3-percentage-point improvement, while retaining the frozen LeWM representation and predictor.

Comment: Criterion 1: LEAP performs action planning through a frozen latent world model by optimizing full action horizons, raising mean success from 77.5% with LeWM+CEM to 94.8% across four control domains.

Relevance: 7 Back to [topic] [top]


2010. The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs

ArXiv: 2609.04110 [page] [pdf]

Authors: Yumeng Shi, Quanyu Long, Yin Wu, Wenya Wang

Abstract: Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token representations where event dynamics should first emerge. This mismatch allows models to learn temporal answers from shortcuts such as objects, scenes, and language priors, without requiring internal video representations to capture event progression. To address this, we propose VT-Contrast, a representation-level temporal counterfactual objective for VideoLMs. Its design asks where temporal supervision should act and what temporal differences it should expose. VT-Contrast supervises selected late-layer last-frame video tokens, where temporal information is expected to be integrated before language generation, and contrasts order-preserving views with same-video reordered counterfactuals graded by Kendall tau distance. It requires no architectural changes, is compatible with diverse VideoLM training tasks, and improves overall performance across temporal understanding benchmarks. Our code is available at https://github.com/ANDgate99/VT-Contrast.

Comment: Matches criterion 2 by introducing VT-Contrast, a representation-level contrastive objective for VideoLM video tokens using order-preserving views versus reordered counterfactuals graded by Kendall tau distance.

Relevance: 8 Back to [topic] [top]

2018. Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision

ArXiv: 2609.04203 [page] [pdf]

Authors: Shravan Venkatraman, Wenshuai Zhao, Mohammad Hassan Vali, Arno Solin

Abstract: We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model generates its own target, so training requires no labels, separate teacher, or reward signal, and adds no inference cost. On LLaVA-OneVision-2-8B, S$^3$T improves VSTAT accuracy by $+1.74$ as a single model, $+2.38$ with souping, and $+2.70$ with additional vision-encoder adaptation, while prior self-evolving methods leave state tracking largely unchanged. The capability learned from unlabeled synthetic clips transfers to real videos, improving performance by $+7.95$ on VSTAT-YouTube state-tracking questions and $+4.50$ on MVBench Action Count.

Comment: Criterion 2: S^3T proposes a self-supervised dense-view teacher/sparse-view student self-distillation method for video state tracking, improving VSTAT-YouTube by +7.95 and MVBench Action Count by +4.50.

Relevance: 7 Back to [topic] [top]


4021. Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

ArXiv: 2609.04183 [page] [pdf]

Authors: Ye-Chan Kim, Seunghee Choi, SeungJu Cha, Si-Woo Kim, Hwiseon Kim, Hyungee Kim, Dong-Jin Kim

Abstract: Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions lack visual grounding and are rigidly assigned to every inter-event gap at a fixed location and duration. To address these, we propose Seeing Before Synthesizing (SBS), a framework that adaptively provides visually grounded linguistic guidance only where warranted. Leveraging a VLM, we generate frame-level narratives for the inter-event gaps and detect transitions from the semantic variation across them. For identified transitions, we then refine inter-event temporal masks by blending the temporal midpoint with the semantic change point and selecting the width that maximizes vision-language alignment. Experiments on ActivityNet Captions and YouCook2 demonstrate state-of-the-art performance in both captioning and localization.

Comment: Criterion 4: SBS leverages a VLM to generate frame-level narratives, detect semantic transition points, and refine temporal masks, achieving state-of-the-art dense video captioning/localization on ActivityNet Captions and YouCook2.

Relevance: 6 Back to [topic] [top]


5002. Sparse auto-regressive modeling for scene generation from multi-view images

ArXiv: 2609.03931 [page] [pdf]

Authors: Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel, Wonjune Cho, Bardienus Pieter Duisterhof, Vincent Leroy, Jerome Revaud

Abstract: Generating complete 3D scenes from sparse, unconstrained views is a fundamental challenge in 3D vision which requires reasoning beyond observed content while remaining computationally tractable. Existing feed-forward reconstruction methods are inherently limited to content visible in the input images, while 3D generative modeling is hindered by the high computational cost of dense volumetric representations and the scarcity of large-scale 3D supervision. We introduce SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion without requiring ground-truth 3D data for supervision. Our key insight is to formulate 3D scene generation in a structured, compact, voxel-aligned 3D latent space where only occupied voxels are represented. We learn this sparse latent space directly from multi-view images using photometric supervision via differentiable 3D Gaussian Splatting. Given a partial set of observed voxels encoded from sparse input views, scene completion reduces to predicting the missing latent tokens and their spatial support within the voxel grid. To this end, we train a masked autoregressive transformer that jointly models voxel occupancy and latent token values, enabling efficient and spatially consistent generation of unseen regions. We demonstrate the effectiveness of our method on synthetic indoor scenes, achieving higher novel-view quality than prior work. We further validate its generalization on RealEstate10k, highlighting its applicability to real-world data.

Comment: Criteria 5 and 6: SPAR3S is a sparse voxel-aligned 3D latent generative model trained from multi-view images via differentiable 3D Gaussian Splatting, using a masked autoregressive transformer and showing higher novel-view quality on synthetic indoor scenes and RealEstate10k.

Relevance: 10 Back to [topic] [top]

5003. VisCAD: A Foundation Model Suite with Multimodal Industrial CAD Intelligence

ArXiv: 2609.03811 [page] [pdf]

Authors: JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong, Zhichao Huang, Guanlin Li, Hongsen Liu, Ziqi Liu, Yichen Long, Luya Wang, Yuchen Wang, Wenxiang Wu, Huimu Yu, Ning Zhang

Abstract: AI-assisted computer-aided design (CAD) for industrial products involves two challenging phases. Part-level generation maps diverse forms of user intent, including renders, text descriptions, 2D drawings, and real photographs, to executable programs in a CAD domain-specific language. Assembly-level generation must additionally handle interacting parts, plan mating relations, estimate poses, and place all parts correctly. Existing specialized CAD models are commonly trained on narrow input domains, such as renders or texts, and often generalize poorly, while general-purpose frontier models cover broader inputs but perform inconsistently across CAD domains. We present VisCAD, a foundation model suite designed to provide both broad generalization and strong CAD capability for realistic industrial products. At its core is VisCAD-M1, a 27B model trained through mid-training and post-training for part-level design generation. On PubCADBench and RealCADBench, VisCAD-M1 achieves the highest average part-level score among the evaluated models, reaching 0.5540 compared with 0.5496 for the strongest frontier model. Reusing VisCAD-M1 as a test-time verifier can further raise the score to 0.5797, an approximately 5 percent relative improvement over the previous state of the art. VisCAD also includes a domain-specific harness that leverages frontier models for complex assembly generation and demonstrates advantages over general-purpose harnesses in both quantitative and qualitative evaluations.

Comment: Criterion 5: VisCAD presents a 27B foundation model suite for image/text/drawing-to-CAD program generation and assembly generation, improving PubCADBench and RealCADBench part-level scores to 0.5540 and 0.5797 with a test-time verifier.

Relevance: 9 Back to [topic] [top]

5009. RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

ArXiv: 2609.03773 [page] [pdf]

Authors: JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong, Zhichao Huang, Guanlin Li, Zongzhen Li, Hongsen Liu, Yichen Long, Wei Wang, Yuchen Wang, Dongyue Yang, Huimu Yu, Xianwen Zhong

Abstract: Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD benchmarks often emphasize synthetic or CAD-native settings, limited input modalities, or executability and IoUs alone. We introduce RealCADBench, a benchmark for intent-to-program CAD modeling from real industrial design intents. It contains 12,632 tasks from 19 factory-automation categories and spans text descriptions, 2D engineering drawings, real product pictures, and rendered images for both Part and Assembly modeling. We report results on a 1,770-task evaluation slice: 1,745 Part tasks across four input regimes and RCB-Assm25, a 25-task assembly study used in every reported assembly comparison. Each method generates FreeCAD API Python, which a shared runtime executes to export the 3D model. We evaluate the exported model using executability, Solid IoU, Surface IoU, and a rubric-based visual-semantic identity Judge. Among the nine standalone frontier large models evaluated, no model leads all four metrics. Across six frontier-scale large models, executability ranges from 0.565 to 0.812, Solid IoU from 0.2841 to 0.5379, and Surface IoU from 0.112 to 0.217 across the four Part regimes. The highest regime-balanced composite comes from a different model than the leaders on the four component metrics. On RCB-Assm25, Codex with GPT-5.5 improves executability and both IoU metrics over standalone GPT-5.5, but lowers the Judge score by 6.98 percentage points, leaving GPT-5.5 as the Judge leader. We also observe recurring failure modes, most notably missing fine structures, loss of part identity, and incorrect assembly placement. These results show that execution alone is insufficient to characterize realistic CAD modeling and that frontier models and agents differ substantially across executability, IoUs, and visual-semantic identity.

Comment: Criterion 5: RealCADBench directly benchmarks multimodal intent-to-program 3D CAD generation from text, 2D drawings, photos, and renders, executing FreeCAD Python and evaluating generated assets with Solid IoU, Surface IoU, and a visual-semantic Judge over 12,632 industrial tasks.

Relevance: 8 Back to [topic] [top]

5011. Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

ArXiv: 2609.04196 [page] [pdf]

Authors: Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy

Abstract: We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.

Comment: Criterion 5: Puffin-World is a generative 3D world model that jointly synthesizes appearance and reconstructs geometry using native physics, depth, image, and Omni-Camera states, scaled with the Puffin-16M vision-language-camera/trajectory dataset.

Relevance: 8 Back to [topic] [top]

5017. Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations

ArXiv: 2609.04174 [page] [pdf]

Authors: Denis M. Akola, David F. Fouhey

Abstract: 3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich unified representations with feed-foward transformers. The scene representations learned by these models enable strong performance on multiple 3D vision tasks. In this paper, we investigate using their internal representations to infer 3D in the scene from new views. Our hypothesis is that in order to solve the task of 3D reconstruction, these models need to learn a representation that includes a large amount of general knowledge about 3D scenes. After showing that it is possible to decode hidden surfaces from internal 3DFM representations, we propose a method, Z3D, that estimates pointmaps in unseen views by doing latent diffusion on 3DFM representation. We show that Z3D can predict realistic depth maps for new views across multiple datasets.

Comment: Matches criterion 5 by using latent diffusion on 3D foundation-model representations in Z3D to synthesize unseen-view pointmaps/depth maps across multiple datasets.

Relevance: 7 Back to [topic] [top]

5020. OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping

ArXiv: 2609.03919 [page] [pdf]

Authors: Zelong Lv, Sicheng Xu, Jianfeng Xiang, Ruicheng Wang, Yue Dong, Yu Deng, Guangzhong Sun, Jiaolong Yang

Abstract: We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-consistent, and high-fidelity visual scenes. Given a single image, OctWorld performs stable autoregressive world generation along user-specified camera trajectories. We focus on long-range generation, characterized by extended camera paths and wide viewpoint coverage, where preserving spatial consistency is particularly challenging when previously generated regions are revisited. To address this problem, we introduce OctMap, an extensible and spatially adaptive 3D memory that progressively fuses generated visual observations and their corresponding depth maps into a global representation. OctMap employs TSDF fusion within a dynamic sparse octree whose spatial resolution adapts to image evidence. This design preserves geometric and appearance details across diverse scene scales while maintaining low memory overhead. Experiments demonstrate that OctWorld generates long-range, spatially consistent videos and outperforms prior methods on both existing benchmarks and challenging long-range generation settings. OctMap also provides clear advantages over point-based caches and fixed-resolution TSDF volumes. Project page: https://maxtirerror.github.io/octworldpage/

Comment: Criterion 5: OctWorld uses a video diffusion model with OctMap, a dynamic sparse-octree TSDF 3D memory, to generate long-range world-consistent explorable scenes from a single image and outperform prior long-range generation methods.

Relevance: 7 Back to [topic] [top]


6004. Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction

ArXiv: 2609.04201 [page] [pdf]

Authors: Chin-Yang Lin, Yang-Che Sun, Cheng Sun, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Wei-Chen Chiu, Yu-Lun Liu

Abstract: Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fixed first-frame anchor forces extrapolation far beyond the training distribution. Small drifts accumulate and amplify into significant geometric collapse. However, we observe that per-frame depth remains stable throughout this failure. The backbone’s local geometry remains intact; only the global pose head breaks down. Motivated by this decoupling, we introduce Scal3R. This approach reformulates online reconstruction as multi-reference relative pose querying. We use lightweight learnable tokens, which make up about ~1% of the parameters, and inject them into a completely frozen backbone via asymmetric attention. This setup queries poses relative to multiple past keyframes. An online pose-graph optimization system with loop closure suppresses long-range drift. Scal3R reaches convergence in 8 hours on a single GPU. It reduces the average ATE by over 60% on KITTI compared to the online baseline. It also achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes. Project page: https://linjohnss.github.io/scal3r/

Comment: Criterion 6: Scal3R performs online 3D reconstruction via multi-reference relative pose querying with lightweight tokens and pose-graph loop closure, reducing average ATE by over 60% on KITTI and reporting SOTA across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes.

Relevance: 9 Back to [topic] [top]

6007. STARS-GS: Structure-Aware Regularized Gaussian Splatting for Large-Scale Aerial Surface Reconstruction

ArXiv: 2609.03447 [page] [pdf]

Authors: Bocheng Li, Wenjuan Zhang, Jie Pan. Dongxu Han, Xuesong Ma, Yiling Yao, Yaning Wang

Abstract: Large-scale 3D surface reconstruction from aerial imagery is fundamental to geospatial mapping and urban modeling. Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated considerable potential for this task. However, existing methods still face three major challenges in large and complex scenes: scene partitioning may split continuous scene elements across independently optimized sub-regions; geometric constraints mainly focus on the attributes of individual Gaussians while overlooking their local organization; and uniform regularization struggles to accommodate heterogeneous geometric structures. To address these issues, we propose STARS-GS, a structure-aware 3DGS framework for large-scale surface reconstruction. First, we introduce a structure-aware scene partitioning strategy that better preserves continuous scene structures during partitioning and reduces cross-region geometric inconsistencies and stitching artifacts through boundary refinement. Second, we develop neighborhood-aware Gaussian organization that extends geometric constraints from individual primitives to their neighborhood organization, encouraging Gaussians to better conform to local surface geometry. Third, we introduce adaptive surface regularization that adjusts the regularization strength according to local geometric characteristics, promoting geometric consistency in structured regions while preserving plausible variations in unstructured regions. Extensive experiments on large-scale aerial photogrammetry benchmarks demonstrate that STARS-GS consistently outperforms the evaluated Gaussian-based methods in surface reconstruction. It increases the average F1-score from 0.640 for the second-best method to 0.698, corresponding to a relative improvement of approximately 9.1\%, demonstrating effective improvements in geometric accuracy and surface completeness.

Comment: Criterion 6: STARS-GS proposes structure-aware 3D Gaussian Splatting for large-scale aerial surface reconstruction, using structure-aware partitioning, neighborhood-aware Gaussian organization, and adaptive surface regularization to improve average F1-score from 0.640 to 0.698 on aerial photogrammetry benchmarks.

Relevance: 8 Back to [topic] [top]

6014. Laplacian Frequency Hierarchies for Efficient 3D Gaussian Splatting Training

ArXiv: 2609.03334 [page] [pdf]

Authors: Yixiong Yang, Sisheng Zhang, Qingsong Yan, Shaohuai Shi, Qiang Wang

Abstract: A key bottleneck in 3D Gaussian Splatting training is the continual growth of Gaussian primitives, which increases optimization cost and slows convergence, especially at high resolutions. We propose Laplacian Frequency Hierarchies, a simple yet efficient 3DGS scheme that combines Laplacian image decomposition with coarse-to-fine, frequency-staged training. After fitting lower-frequency structure, we archive the corresponding Gaussian field so that subsequent fields can optimize higher-frequency residuals without carrying the full primitive burden, and we compose the rendered components in the image domain via a Laplacian-style reconstruction at inference time. This design reduces the number of active Gaussians during training, thereby lowering optimization overhead and accelerating training. The proposed scheme is plug-and-play and orthogonal to prior 3DGS accelerations: it can be directly combined with strong backbones such as Taming-3DGS and FastGS to improve training speed with competitive reconstruction quality. It achieves average speedups of 1.73x and 1.21x at 1K setting, and 1.74x and 1.33x at 4K setting on Taming-3DGS and FastGS, with larger gains on more challenging scenes and increasingly pronounced benefits at higher resolutions.

Comment: Criterion 6: proposes Laplacian Frequency Hierarchies for 3D Gaussian Splatting training, achieving 1.73x/1.74x speedups at 1K/4K on Taming-3DGS with competitive reconstruction quality.

Relevance: 7 Back to [topic] [top]

6015. TruncGradGS: Improved 3D Gaussian Splatting via Truncated Gradient Updates

ArXiv: 2609.03534 [page] [pdf]

Authors: Theo Morales, Nhat-Quynh Le-Pham, Robin Atkins, Binh-Son Hua

Abstract: 3D Gaussian Splatting has become a de facto scene representation for novel view synthesis, yet robustly learning 3D Gaussian primitives from visual input remains challenging. Standard optimization relies on gradient-based updates, but a common issue is the gradient vanishing phenomenon: a pixel far from a Gaussian primitive often has diminishing gradient magnitudes to influence primitive attributes, resulting in suboptimal scene reconstruction. In this paper, we propose a method to address gradient vanishing with a piecewise truncated gradient formulation that improves the optimization stability and robustness to initializations. We show that our method consistently improves 3D Gaussian Splatting with random and COLMAP initializations while being generalizable across static and dynamic Gaussian Splatting. As a by-product, we also examine the limitations of current benchmarks for dynamic scenes, and introduce a novel dataset for benchmarking dynamic Gaussian Splatting using synthetic 3D scenes. We demonstrate the effectiveness of our method in both static and dynamic settings for the public benchmarks and our proposed dataset.

Comment: Criterion 6: introduces a piecewise truncated-gradient update for 3D Gaussian Splatting that improves static and dynamic reconstruction robustness under random and COLMAP initializations.

Relevance: 7 Back to [topic] [top]

6016. Stable and Scalable Bundle Adjustment of Holistic 3D Structures

ArXiv: 2609.04026 [page] [pdf]

Authors: Shaohui Liu, R'emi Pautrat, Daniel Barath, Richard Hartley, Viktor Larsson, Marc Pollefeys

Abstract: Bundle Adjustment (BA) is a cornerstone of 3D computer vision and has benefited from decades of advances in sparse optimization and numerical methods. It was originally developed for jointly optimizing camera intrinsics, poses and sparse 3D points. While extensions incorporate lines and other primitives, integrating richer geometric structures such as parallelism, coplanarity, or wireframes often introduces significantly increased computational cost and reduced numerical stability. In this paper, we propose a unified framework that extends bundle adjustment to jointly optimize geometric features and higher-order relations. We first introduce a taxonomy that distinguishes scalable geometric features with direct 2D measurements (e.g., points and lines), from groups encoding higher-order relations (e.g., coplanarity, parallelism, etc.), where we show that groups can be modeled as camera-like entities within the bundle adjustment framework. Building on this formulation, we propose that both group constraints and cross-feature relations (i.e., point-line associations) can be expressed through 2D reprojection measurements. By formulating group-induced and cross-feature reprojection errors, we preserve the sparsity structure of classical point-based BA under Schur elimination, while avoiding direct 3D regularization that degrades the conditioning and stability. Experiments on both real-world and synthetic datasets demonstrate runtime performance comparable to classical point-only bundle adjustment, while producing significantly richer 3D structures and improved geometric accuracy.

Comment: Matches criterion 6 by proposing a unified bundle-adjustment framework that jointly optimizes camera intrinsics/poses with points, lines, and higher-order structures such as coplanarity and parallelism while preserving Schur sparsity.

Relevance: 7 Back to [topic] [top]

6022. VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues

ArXiv: 2609.03824 [page] [pdf]

Authors: Ernesto Lozano, Alberto Jaenal, Javier Civera

Abstract: 3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene, showcasing strong zero-shot generalization. However, as metric scale is not observable from monocular images, their absolute scale predictions are typically inaccurate. Inertial measurement units (IMUs), present in most devices, naturally complement monocular cameras by observing scaled motion. We introduce VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings. VI3 initializes and preintegrates the IMU to obtain a metric motion reference, which is then used to recover the scale of the 3DFM outputs. Our method includes adaptable anchoring strategies tailored to diverse 3DFM architectures. Experiments on synthetic and real aerial datasets demonstrate that VI3 recovers metric scale without ground-truth supervision while preserving geometric consistency, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.

Comment: Criterion 6: VI3 improves multi-view 3D reconstruction by metrically anchoring pretrained 3D foundation-model pose/depth outputs using IMU preintegration, recovering scale without ground-truth supervision on synthetic and real aerial datasets.

Relevance: 6 Back to [topic] [top]