Publications

publications by categories in reversed chronological order. generated by jekyll-scholar.

2026

  1. arXiv
    cao2026xiaomigui.png
    Xiaomi-GUI-0 Technical Report
    Wanxia Cao, Chengzhen Duan, Pei Fu, and 29 more authors
    arXiv preprint arXiv:2606.31410, 2026
  2. ECCV
    lyu2026unitranslator.png
    UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation
    J Lyu, P Fu, Z Li, and 6 more authors
    In European Conference on Computer Vision (ECCV), 2026
  3. arXiv
    sun2026beyond.png
    Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment
    Y Sun, P Fu, Shaojie Zhang, and 6 more authors
    arXiv preprint arXiv:2605.14311, 2026
  4. arXiv
    xu2026qmask.png
    Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models
    L Xu, F Feng, Shaojie Zhang, and 7 more authors
    arXiv preprint arXiv:2604.00161, 2026
  5. arXiv
    lyu2026imtbench.png
    IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation
    J Lyu, P Fu, Z Li, and 7 more authors
    arXiv preprint arXiv:2603.10495, 2026
  6. ECCV
    wang2026gaia.png
    GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
    S Wang, P Fu, R Zhang, and 6 more authors
    In European Conference on Computer Vision (ECCV), 2026

2025

  1. ICCV
    zhang2025qframe.png
    Q-Frame: Query-aware Frame Selection and Multi-resolution Adaptation for Video-LLMs
    Shaojie Zhang, J Yang, Jianqin Yin, and 2 more authors
    In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025
  2. NeurIPS
    zhang2025btlui.png
    BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
    Shaojie Zhang, R Zhang, P Fu, and 6 more authors
    In Advances in Neural Information Processing Systems (NeurIPS), 2025
  3. arXiv
    zhang2025hyperclick.png
    HyperClick: Advancing Reliable GUI Grounding via Uncertainty Calibration
    Shaojie Zhang, P Fu, R Zhang, and 6 more authors
    arXiv preprint arXiv:2510.27266, 2025
  4. PR
    zhang2025contrastive.png
    A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
    Shaojie Zhang, Jianqin Yin, and Yonghao Dang
    ELSEVIER Pattern Recognition (PR), 2025
  5. IROS
    chen2025aligncape.png
    AlignCAPE: Support and Query Feature Aligning for Category-Agnostic Pose Estimation
    Z Chen, J Tang, G Xu, and 3 more authors
    In 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025
  6. IROS
    wei2025masksem.png
    MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
    W Wei, Shaojie Zhang, Yonghao Dang, and 1 more author
    In 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025

2024

  1. ROBIO
    liu2024temporal.png
    Temporal Text Prompts for Skeleton-based Action Recognition
    Liyuan Liu, Shaojie Zhang, Yonghao Dang, and 2 more authors
    In 2024 IEEE International Conference on Robotics and Biomimetics (ROBIO), 2024
  2. TCSVT
    zhang2024sitmlp.png
    SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
    Shaojie Zhang, Jianqin Yin, Yonghao Dang, and 1 more author
    IEEE Transactions on Circuits and Systems for Video Technology (TCSVT), 2024
  3. KBS
    xu2024mlpair.png
    MLP-AIR: An Effective MLP-based Module for Actor Interaction Relation Learning in Group Activity Recognition
    Guoliang Xu, Jianqin Yin, Shaojie Zhang, and 1 more author
    ELSEVIER Knowledge-Based Systems (KBS), 2024
  4. PR
    dang2024kinematics.png
    Kinematics Modeling Network for Video-based Human Pose Estimation
    Yonghao Dang, Jianqin Yin, Shaojie Zhang, and 2 more authors
    ELSEVIER Pattern Recognition (PR), 2024

2023

  1. arXiv
    fu2023improved.jpg
    An Improved Baseline Framework for Pose Estimation Challenge at ECCV 2022 Visual Perception for Navigation in Human Environments Workshop
    J Fu, Yonghao Dang, R Yin, and 4 more authors
    arXiv preprint arXiv:2303.07141, 2023

2022

  1. TIP
    dang2022relation.png
    Relation-Based Associative Joint Location for Human Pose Estimation in Videos
    Yonghao Dang, Jianqin Yin, and Shaojie Zhang
    IEEE Transactions on Image Processing (TIP), 2022