Jiuyi Xu

Jiuyi (Joey) Xu徐久乙

Ph.D. Student in Robotics

Colorado School of Mines, Golden CO · Advised by Dr. Yangming Shi

Efficient AI Robot Learning Vision-Language-Action (VLA) World-Action Model (WAM)
About

Jiuyi (Joey) Xu is a Ph.D. student in Robotics at the Colorado School of Mines, advised by Dr. Yangming Shi. He works at the crossroads of efficient AI and embodied intelligence, aiming to make large vision-language-action (VLA) and world-action models (WAMs) fast and deployable on real robots with consumer-grade GPU. He is currently a Robotics Intern at Prime Robotics, building real-time multi-agent path planning and LLM-driven robot agents.

Earlier, he earned an M.S. in Computer Science at USC and a B.E. in Software Engineering at the Dalian University of Technology, and interned at USC's Institute for Creative Technologies on open-vocabulary perception and 3D reconstruction. He reviews for venues including NeurIPS, AISTATS, and IEEE TNNLS. Away from the desk, he is usually at the gym or on the basketball court.

Background

Education & Experience

Education

Ph.D. in Robotics2024 – Present

Colorado School of Mines, Golden, CO · Advisor: Dr. Yangming Shi

M.S. in Computer Science2022 – 2024

University of Southern California, Los Angeles, CA

B.E. in Software Engineering2017 – 2021

Dalian University of Technology, Dalian, China

Research Experience

Graduate Research Assistant — Shi Research GroupAug 2024 – Present

Colorado School of Mines, Golden, CO · Advisor: Dr. Yangming Shi

  • Develop efficient vision-language-action (VLA) models and world-action models (WAMs) for embodied agents.
  • Design diffusion-based generative pipelines for egocentric image synthesis on construction sites (EgoVision, CRC 2026).
  • Apply synthetic image generation to construction safety monitoring (GenVision, WSC 2025).
Research Assistant (Intern) — Geospatial Terrain Research, USC-ICTJun 2023 – Jun 2024

USC Institute for Creative Technologies, Los Angeles, CA · Advisor: Dr. Meida Chen

  • Automated annotation of real and synthetic imagery via open-vocabulary object detection (OVOD) and semantic segmentation.
  • Built a tripping-hazard detection pipeline for construction sites using Grounding DINO and the Segment Anything Model (SAM).
  • Developed 2D–3D open-vocabulary semantic segmentation (OVSS) for photogrammetric point clouds via OVOD and k-means clustering.

Industry Experience

Robotics Intern — Prime RoboticsMay 2026 – Present

Denver, CO · Manager: Ryan Austin

  • Develop and maintain core components of a real-time multi-agent path-planning system coordinating 10+ robots in production.
  • Build LLM-driven robot agents via the Model Context Protocol (MCP), enabling robots to reason about and execute tasks from their own embodied perspective.
Research

Publications

2026

VLAQuantBench: Benchmarking Post-Training Quantization for Vision-Language-Action Models under Closed-Loop Control
Xu, J., Chen, M., Wang, S., Sui, Y., Shi, Y.
Preprint
FastVLA: Instruction-Conditioned and Temporally Adaptive Visual Token Compression for Efficient Vision-Language-Action Models
Xu, J., Qian, C., Tu, Z., Sui, Y., Shi, Y.
Preprint
Domain Knowledge-Infused Prompting for Deep and Full-Sentence Information Extraction from Bridge Inspection Reports
Vahidi, M., Xu, J., Shi, Y., Liu, K.
Journal of Computing in Civil Engineering
An Autoregressive-based Framework for Simulating Stochastic Occupant Behavior in Residential Buildings
Xu, J., Kim, R., Ye, Y., Shi, Y.
Building SimulationSpringer ↗
GenVision: Enhancing Construction Safety Monitoring with Synthetic Image Generation
Xu, J., Chen, M., Shi, Y.
WSC 2025IEEE ↗

2025

LowDiff: Efficient Diffusion Sampling with Low-Resolution Condition
Xu, J., Jin, Q., Chen, M., Feng, A., Sui, Y., Shi, Y.
PreprintarXiv ↗
IDU: Incremental Dynamic Update of Existing 3D Virtual Environments with New Imagery Data
Chen, M., Leal, L., Hu, Y., Liu, R., Xiong, B., Feng, A., Xu, J., Shi, Y.
I/ITSEC 2025arXiv ↗
Egocentric Camera-Based Method for Detecting Static Hazardous Objects on Construction Sites
Liu, Z., Xu, J., Wun, C., Chen, M., Zou, Z., Shi, Y.
Automation in ConstructionDOI ↗

2024

Open-Vocabulary High-Resolution 3D (OVHR3D) Data Segmentation and Annotation Framework
Xu, J., Chen, M., Feng, A., Yu, Z., Shi, Y.
I/ITSEC 2024arXiv ↗
To Shelter or Not To Shelter: Exploring the Influence of Different Modalities in Virtual Reality on Individuals' Tornado Mitigation Behaviors
Xu, J., Sanni, T., Liu, Z., Yang, Y., Lee, J., Song, W., Shi, Y.
PreprintarXiv ↗
Large-Scale 3D Terrain Reconstruction Using 3D Gaussian Splatting for Visualization and Simulation
Chen, M., Lal, D., Yu, Z., Xu, J., Feng, A., You, S., Nurunnabi, A., Shi, Y.
ISPRS ArchivesDOI ↗
An Open-Vocabulary Framework for Efficient 2D and 3D Visual Data Annotation in Indoor and Outdoor Built Environments
Rahman, M.A., Xu, J., Liu, Z., Chen, M., Zhu, R., Liu, Y., Shi, Y., Feng, A.
i3ce 2024
Detecting without Training: An Open-Vocabulary Object Detection Method for Identifying Hazardous Objects on Construction Sites
Liu, Z., Xu, J., Suen, W.K., Chen, M., Zou, Z., Feng, A., Shi, Y.
i3ce 2024
Selected Work

Projects

LowDiff overviewLowDiff

LowDiff

DiffusionEfficient SamplingImage Generation

A cascaded diffusion framework that generates increasingly higher-resolution outputs, but replaces the usual stack of stage-specific models with a single unified model that progressively refines images from low to target resolution. It reaches comparable or better quality with far fewer high-resolution sampling steps and applies to both pixel-space and latent-space diffusion. Across CIFAR-10, FFHQ and ImageNet (conditional and unconditional) it delivers 50–80% throughput gains — e.g. FID 1.94 on conditional CIFAR-10 and FID 1.59 on ImageNet-256 (LightningDiT-XL/1) — while maintaining quality.

FastVLA overviewFastVLA

FastVLA

VLAToken CompressionTraining-Free

A training-free, instruction-aware, and temporally adaptive visual-token compression framework for efficient closed-loop VLA inference. At each step it scores token importance via instruction-to-vision attention and sets the compression rate from the importance-distribution uniformity, measured by an inverse-Simpson metric with online sliding-window calibration. Across OpenVLA, OpenVLA-OFT and CogACT on LIBERO and SIMPLER it reduces LLM-backbone FLOPs by 35–59% and action latency by 9–32% (1.09×–1.46× speedup) while maintaining task success; on a physical Franka Research 3 robot it cuts end-to-end latency by 20% (1.25× speedup) at comparable success.

QuantVLA main results — post-training quantization across VLA models and benchmarksQuantVLA

QuantVLA

VLAQuantizationPTQ Benchmark

A benchmark that evaluates post-training quantization (PTQ) of VLA models entirely under closed-loop task success, arguing that action error alone is an insufficient proxy for deployment. It spans 7 VLA models, 4 simulation environments and 6 weight / weight-activation bit settings, covering both full-model quantization and fine-grained, component-wise sensitivity across the LLM backbone, vision encoder and action head. It further benchmarks off-the-shelf LLM quantization methods on the VLA backbone, measures inference latency and memory under real quantization kernels, and validates behavior on a physical robot.

Recognition

Honors & Service

Awards

Academic Service — Reviewer

NeurIPS 2026 AISTATS 2026 IEEE TNNLS Int'l Journal of Data Science & Analytics J. Computing in Civil Engineering ISARC 2026 ISARC 2025
Updates

News

Get in touch

Contact

Happy to chat about VLAs, WAMs, efficient AI, or potential collaborations.