Reframing 3D Point-Cloud Learning And Road-Scene Geometry: Measurement Chains and Validation Design: Hierarchical Spatial Mamba Framework and WaveSamba Wavelet Transform SSM
PDF

Keywords

3D Point-Cloud Learning And Road-Scene Geometry
Neighborhood Construction
Hierarchy
State-Space Mixing
Sampling Robustness
Efficiency

Abstract

The literature on 3D point-cloud learning and road-scene geometry contains a recurring tension between methodological novelty and evidential comparability. By reading hierarchical spatial state-space aggregation for point-cloud classification alongside a wavelet-transform state-space decoder for zero-shot depth estimation, this article clarifies the conditions under which their conclusions can support a common research argument. The analysis combines two focal publications with 12 previously verified sources and organizes the evidence around neighborhood construction, hierarchy, state-space mixing, sampling robustness, and efficiency. Rather than pooling incompatible outcomes, it compares research questions, representations, controls, and validation envelopes. Comparison reveals recurring trade-offs among neighborhood construction, hierarchy, and state-space mixing. These trade-offs do not support a universal ranking; instead, they identify the operating envelope within which each method remains credible and the perturbations most likely to expose fragile conclusions. The article concludes with a research agenda built around transparent comparators, targeted stress tests, and evidence records that can be reused without overstating causal or practical reach.

PDF

References

Sun, Y., Zia, A., Long, Z., Qiu, Z., Xiang, W., & Zhou, J. (2025). Hierarchical Spatial Mamba Framework for Point Cloud Classification. In Pattern Recognition and Computer Vision (pp. 402-417). Springer.

Hou, B., & Long, Z. (2024). WaveSamba: A Wavelet Transform SSM Zero-Shot Depth Estimation Decoder. In 2024 International Conference on Digital Image Computing: Techniques and Applications (DICTA) (pp. 738-744). IEEE.

Zhang, T., Yuan, H., Qi, L., Zhang, J., Zhou, Q., Ji, S., et al. (2025). Point Cloud Mamba: Point Cloud Learning via State Space Model. Proceedings of the AAAI Conference on Artificial Intelligence, 39(10), 10121-10130. https://doi.org/10.1609/aaai.v39i10.33098

Tan, J., Li, J., An, X., & He, H. (2014). Robust Curb Detection with Fusion of 3D-Lidar and Camera Data. Sensors, 14(5), 9046-9073. https://doi.org/10.3390/s140509046

Song, S., Tang, K., & Zhang, Y. (2025). PST-Mamba: Spatio-temporal selective state fusion for effective point cloud video understanding with state space models. Image and Vision Computing, 163, 105785. https://doi.org/10.1016/j.imavis.2025.105785

Liu, J., Yue, S., Hao, W., & Cai, Y. (2026). MF-BEVFusion: multiscale depth estimation and fully dynamic fusion for camera-LiDAR BEV 3D object detection. Journal of Electronic Imaging, 35(02). https://doi.org/10.1117/1.jei.35.2.023009

Zhou, Z., Wang, Q., & Zhou, X. (2026). MSHI-Mamba: A Multi-Stage Hierarchical Interaction Model for 3D Point Clouds Based on Mamba. Applied Sciences, 16(3), 1189. https://doi.org/10.3390/app16031189

Bong, E. J., & Kee, S. C. (2026). Dense Depth Map Estimation Based on Camera–LiDAR Sensor Fusion. IEEE Sensors Journal, 26(4), 5891-5901. https://doi.org/10.1109/jsen.2025.3649237

Wang, G., Zhang, X., Peng, Z., Zhang, T., & Jiao, L. (2025). S 2 Mamba: A Spatial–Spectral State Space Model for Hyperspectral Image Classification. IEEE Transactions on Geoscience and Remote Sensing, 63, 1-13. https://doi.org/10.1109/tgrs.2025.3530993

Rao, R., Ouyang, Z., Chen, S., Chen, L., Huang, G., & Cui, C. (2026). Zero-Shot Polarization-Intensity Physical Fusion Monocular Depth Estimation for High Dynamic Range Scenes. Photonics, 13(3), 268. https://doi.org/10.3390/photonics13030268

Liao, J., & Wang, L. (2026). SSA-Mamba: Spatial-Spectral Attentive State Space Model for Hyperspectral Image Classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 19, 6403-6424. https://doi.org/10.1109/jstars.2026.3654346

Ji, M., Yang, J., & Zhang, S. (2026). DepthFusion: Depth-Aware Hybrid Feature Fusion for LiDAR-Camera 3D Object Detection. IEEE Transactions on Multimedia, 28, 7217-7227. https://doi.org/10.1109/tmm.2026.3668596

Xi, G., Wang, C., Liu, X., Xiao, B., & Wei, X. (2026). Sparse Point Cloud Classification Method Based on MSE-Mamba. Electronics, 15(14), 3087. https://doi.org/10.3390/electronics15143087

Obando-Ceron, J. S., Romero-Cano, V., & Monteiro, S. (2023). Probabilistic multi-modal depth estimation based on camera–LiDAR sensor fusion. Machine Vision and Applications, 34(5). https://doi.org/10.1007/s00138-023-01426-x

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Copyright (c) 2026 Spencer Fowler, Tanner Mercer, Troy Benson (Author)