Assistive Visual Navigation: Intelligent Edge Perception & Closed-Loop Guidance
2026
PI: Dr. Abolfazl Razi [arazi@clemson.edu]
AI-SENDS Lab — School of Computing, Clemson University
In collaboration with the University of South Carolina, Arizona State University, SC Citadel, South Carolina State University, and USC Beaufort
Mission & Vision
This project develops AI-powered closed-loop assistive navigation tools to support blind and visually impaired individuals during dynamic locomotion, ranging from daily walking to active travel and running. To enable safe, independent mobility, we are conducting research across multiple dimensions: combining ego-motion estimation and denoising (compensating for natural head motion, body sway, and camera jitter without manual calibration to track true locomotion intent at over 40 FPS) with autonomous user direction sensing (motor focus) to disentangle camera orientation from actual physical movement; performing real-time scene understanding and hazard interpretation via open-vocabulary detectors and depth-conditioned vision-language models (VLMs); translating spatial perception into intuitive, low-latency cadence-adaptive spatial acoustic cues and haptic feedback; and augmenting local edge control by distilling foundation multimodal reasoning (VLMs/VLAs) into compact models executing on mobile neural engines for offline reliability under 60 ms latency.
Sponsored by the South Carolina EPSCoR Program (Award # 26-CRP03), this initiative brings together a collaborative team of researchers across Clemson University and participating organizations (University of South Carolina, Arizona State University, USC Beaufort, South Carolina State University, and The Citadel) to develop methods, tools, and datasets as follows!

Live System Demonstration: Real-time egocentric perception and conversational VLM guidance delivered to a visually impaired user during dynamic navigation.
Interactive System Demonstrations
![]() Open-World Tracking | ![]() Ego-Motion Compensation | ![]() Acoustic Spatial Feedback |
Core Research Pillars
Pillar A: Cadence-Adaptive Closed-Loop Guidance
Conventional electronic travel aids rely on fixed controller parameters optimized for a single walking speed, causing severe latency and oscillatory “Z-walk” overcorrections when users accelerate.
- Gait Rhythm Coupling: By capturing real-time device motion at 100 Hz from a torso/waist-mounted sensor, our controller estimates cadence without requiring user-specific stride calibration.
- Dynamic Control Law: An affine adaptation law continuously adjusts look-ahead preview horizons, temporal smoothing, and derivative damping in real time.
- Empirical Validation: In controlled locomotion trials, cadence-adaptive control reduced corrective cue reversals by 71% ($p=0.001$), decreased correction latency by 52%, and suppressed large guideline departures by 48% while maintaining a steady 24 FPS throughput.

Pillar B: Edge-Computed Spatial Reasoning & Depth Prior Fusion
Large vision-language models (VLMs) frequently suffer from hallucination and temporal drift across long-horizon egocentric videos.
- Metric Depth Conditioning: We integrate Depth Anything 3 (DA3) metric depth maps directly into RGB streams, injecting geometric inductive bias without altering underlying model weights.
- Compact Model Generalization: Benchmarks demonstrate that compact 0.5B–2B parameter models (e.g., InternVL2-2B, SmolVLM2) generalize more effectively to spatial obstruction detection than ungrounded 7B/8B counterparts, enabling full on-device execution.
- Impairment Simulation: To ensure clinical relevance, perception algorithms are evaluated under simulated visual pathologies, including retinitis pigmentosa (tunnel vision), macular degeneration, and cataract scatter.

Pillar C: All-Pixel Ego-Motion & Hazard Prioritization
Body-worn and handheld cameras suffer from extreme high-frequency shaking that degrades conventional visual odometry.
- Calibration-Free Ego-Motion Compensation: Rather than relying on computationally heavy sparse keypoint matching (SIFT/ORB), our approach treats every pixel as a flow vector and solves rigid frame transformations using closed-form Singular Value Decomposition (SVD) in under 1 ms.
- Spatial H-Partitioning: Frames are dynamically segmented into four functional safety zones (Ground, Front, Left, Right) to differentiate between background clutter and actionable hazards requiring immediate deceleration.
- Edge Hardware Benchmarking: The end-to-end perception pipeline achieves 51–61 ms inference on mobile neural engines (Apple A16/M2 architectures).

Benchmark Datasets & Resources
We contribute open datasets and evaluation protocols to advance embodied AI for assistive mobility:
- VIN-Bench (Visual Assistive Navigation Benchmark):
A large-scale multimodal corpus comprising 10,046 event-anchored clips across 70 sessions (>50 hours). Synchronizes dual-view egocentric RGB (head-mounted POV vs. chest-mounted primary), 100 Hz IMU, LiDAR depth, and GPS trajectories with sub-100 ms cross-device acoustic alignment. Supports standardized evaluation of Vision-Language-Action (VLA) trajectory planning and safety reasoning. - Sanpo-D Spatial VQA Dataset:
A navigation-oriented spatial re-annotation of the Google Sanpo dataset, featuring 647 fine-grained multimodal QA pairs probing pedestrian proximity, lateral obstructions, intersection layouts, and path affordances under long-horizon drift.
👉 Dataset Access: Hugging Face Sanpo-D Repository - Urban Cruising & Motor Focus Corpus:
Real-world first-person walking, biking, and scooter recordings annotated for frame-by-frame motor focus intention and ego-motion compensation.
👉 Dataset Access: Google Drive Archive
Research Products
- Cadence-Adaptive Control for Camera-Based Assistive-Running Guidance
Si-En Hong, Hao Wang, Abolfazl Razi
IEEE International Conference on Wearable and Implantable Body Sensor Networks (BSN), 2026.
[Accepted]📖 Show BibTeX
@inproceedings{hong2026cadence, title={Cadence-Adaptive Control for Camera-Based Assistive-Running Guidance}, author={Hong, Si-En and Wang, Hao and Razi, Abolfazl}, booktitle={IEEE International Conference on Wearable and Implantable Body Sensor Networks (BSN)}, year={2026} } - Spatial-Conditioned Reasoning in Long-Horizon Egocentric Videos
James Tribble, Si-En Hong, Hao Wang, Chaoyi Zhou, Ashish Bastola, Siyu Huang, Abolfazl Razi
IEEE International Conference on Wearable and Implantable Body Sensor Networks (BSN), 2026.
[Accepted]• Dataset (HuggingFace)📖 Show BibTeX
@inproceedings{tribble2026spatial, title={Spatial-Conditioned Reasoning in Long-Horizon Egocentric Videos}, author={Tribble, James and Hong, Si-En and Wang, Hao and Zhou, Chaoyi and Bastola, Ashish and Huang, Siyu and Razi, Abolfazl}, booktitle={IEEE International Conference on Wearable and Implantable Body Sensor Networks (BSN)}, year={2026} } - Motion Focus Recognition in Fast-Moving Egocentric Video
Si-En Hong, James Tribble, Alexander Lake, Hao Wang, Chaoyi Zhou, Ashish Bastola, Siyu Huang, Eisa Chaudhary, Brian Canada, Ismahan Arslan-Ari, Abolfazl Razi
IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVw @ CV4WS), 2026.
arXiv:2601.07154📖 Show BibTeX
@article{hong2026motion, title={Motion Focus Recognition in Fast-Moving Egocentric Video}, author={Hong, Si-En and Tribble, James and Lake, Alexander and Wang, Hao and Zhou, Chaoyi and Bastola, Ashish and Huang, Siyu and Chaudhary, Eisa and Canada, Brian and Arslan-Ari, Ismahan and Razi, Abolfazl}, journal={arXiv preprint arXiv:2601.07154}, year={2026} } - VIN-Bench: Benchmarking Safety Reasoning and Action Planning for Visual Assistive Navigation
IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2027 Datasets Track.
[Under Review]📖 Show BibTeX
@article{vinbench2027, title={VIN-Bench: Benchmarking Safety Reasoning and Action Planning for Visual Assistive Navigation}, author={Anonymous}, journal={WACV Datasets Track (Under Review)}, year={2027} } - Motor Focus: Fast Ego-Motion Prediction for Assistive Visual Navigation
Hao Wang, Jiayou Qin, Xiwen Chen, Ashish Bastola, John Suchanek, Zihao Gong, Abolfazl Razi
IEEE 20th International Conference on Body Sensor Networks (BSN), 2024.
DOI: 10.1109/BSN63547.2024.10780583 • Poster📖 Show BibTeX
@inproceedings{wang2024motor, title={Motor Focus: Fast Ego-Motion Prediction for Assistive Visual Navigation}, author={Wang, Hao and Qin, Jiayou and Chen, Xiwen and Bastola, Ashish and Suchanek, John and Gong, Zihao and Razi, Abolfazl}, booktitle={2024 IEEE 20th International Conference on Body Sensor Networks (BSN)}, pages={1--4}, year={2024}, organization={IEEE} } - VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
Hao Wang, Jiayou Qin, Ashish Bastola, Xiwen Chen, John Suchanek, Zihao Gong, Abolfazl Razi
arXiv preprint arXiv:2403.12415, 2024.
arXiv:2403.12415 •50+ Citations📖 Show BibTeX
@article{wang2024visiongpt, title={VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation}, author={Wang, Hao and Qin, Jiayou and Bastola, Ashish and Chen, Xiwen and Suchanek, John and Gong, Zihao and Razi, Abolfazl}, journal={arXiv preprint arXiv:2403.12415}, year={2024} }Project Team
- Abolfazl Razi — Clemson University
- Daniel Si-En Hong — Clemson University
- Hao Wang — Arizona State University
- Ashish Bastola — Clemson University
- Dr. Ismahan Arslan-Ari — University of South Carolina
Special Thanks to Our Collaborators & Contributors
- Dr. Brian Canada — University of South Carolina Beaufort
- Dr. Nikunja K. Swain — South Carolina State University
- Dr. Farhath Zareen — The Citadel, The Military College of South Carolina
- Dr. Zihao Gong — Tokai University
- James Tribble — Clemson University
- Xiwen Chen — Stanley Morgan Company
- Jiayou Qin — Stevens Institute of Technology
- Chaoyi Zhou — Clemson University
- Alexander Lake — Clemson University
- Siyu Huang — Clemson University
- Eisa Chaudhary — University of South Carolina Beaufort
- John Suchanek — Clemson University
Acknowledgments
Research reported in this project is supported in part by the South Carolina EPSCoR Program under award number SC EPSCoR 26-CRP03 and the National Science Foundation (NSF) under award number NSF Award # OIA-2242812. The views, perspectives, and content do not necessarily represent the official views of the SC EPSCoR Program nor those of the NSF.



