Home  /  Research Papers  /  Surveys

*: Equal Contribution, †: Corresponding Author.

Research Papers

2026

Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
Jianmin Chen*, Jiaqi Tang*, Wei Wei†, Xiaogang Xu, Jiafei Wu, Zhe Liu, Qianzhou Wang, Yingying Yan, Botong Geng, Yuyang Xia, Lei Zhang, Qifeng Chen
ACM International Conference on Multimedia (ACM MM), 2026
arXiv / code / model / data / bibtex

Media Report: AI Era (新智元)

Process-level rewards that sustain visual grounding in long reasoning chains.

IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
Jinjian Wu*, Jiaqi Tang*, Wei Wei†, Yingying Yan, Jianmin Chen, Botong Geng, Lei Zhang, Qifeng Chen†
European Conference on Computer Vision (ECCV), 2026
arXiv / code / model / data / demo / bibtex

Media Report: QbitAI (量子位)

Image quality assessment grounded in tool-generated perceptual evidence.

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?
Jiaqi Tang*, Jianmin Chen*, Youyang Zhai*, Wei Wei†, Runtao Liu, Mengjie Zhao, Xiangyu Wu, Qingfa Xiao, Qifeng Chen†
International Conference on Machine Learning (ICML), 2026
arXiv / code / model / demo / bibtex

Media Report: QbitAI (量子位) / HuggingFace Daily Papers

MLLMs that self-recover corrupted inputs before reasoning over them.

LongVideoAgent: Multi-Agent Reasoning with Long Videos
Runtao Liu, Ziyi Liu, Jiaqi Tang, Yue Ma, Renjie Pi, Jipeng Zhang, Qifeng Chen
Annual Meeting of the Association for Computational Linguistics (ACL), 2026
arXiv / code / project page / bibtex

Media Report: HuggingFace Daily Papers

Master agent coordinates grounding and vision agents over hour-long videos.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding
Ke Ma*, Jiaqi Tang*, Bin Guo, Xueting Han, Ruonan Xu, Qingfeng He, Ziheng Wang, Xu Wang, Qifeng Chen, Zhiwen Yu, Yunhao Liu
Annual Meeting of the Association for Computational Linguistics (ACL), 2026
arXiv / code / bibtex

Media Report: Synced (机器之心)

Scene graphs align streaming evidence with query conditions to time responses.

LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
Jiaqi Tang*, Yu Xia*, Yi-Feng Wu*, Yuwei Hu*, Yuhui Chen, Qing-Guo Chen, Xiaogang Xu†, Xiangyu Wu, Hao Lu, Yanqing Ma, Shiyin Lu, Qifeng Chen†
Findings of the Association for Computational Linguistics (ACL Findings), 2026
arXiv / code / bibtex

Entropy-guided targets and a distance-aware reward sharpen GUI grounding.

ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall
Jiayu Yang, Yuxuan Fan, Songning Lai, Shengen Wu, Jiaqi Tang, Chun Kang, Zhijiang Guo, Yutao Yue†
International Conference on Learning Representations (ICLR), 2026
arXiv / code / bibtex

Editing the query-value neuron pathways that carry multi-hop knowledge.

Adaptive Debiasing Tsallis Entropy for Test-Time Adaptation
Xiangyu Wu, Dongming Jiang, Yueying Tian, Feng Yu, Qing-Guo Chen, Jiaqi Tang, Yang Yang, Jianfeng Lu†
International Conference on Learning Representations (ICLR), 2026
code / bibtex

Class-adaptive Tsallis entropy that debiases CLIP at test time.

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
Jiaqi Tang*, Jianmin Chen*, Wei Wei†, Xiaogang Xu, Runtao Liu, Xiangyu Wu, Qipeng Xie, Jiafei Wu, Lei Zhang, Qifeng Chen†
AAAI Conference on Artificial Intelligence (AAAI), 2026   (Oral)
arXiv / code / model / data / demo / talk / bibtex

Media Report: AI Era (新智元) / Synced (机器之心) / CVer / HuggingFace Daily Papers / VALSE 2026

Explicit degradation reasoning with intensity-adaptive depth.

Sage Deer: A Super-Aligned Driving Generalist Is Your Copilot
Hao Lu*, Jiaqi Tang*, Jiyao Wang, Yunfan LU, Xu Cao, Qingyong Hu, Yin Wang, Yuting Zhang, Tianxin Xie, Yunpeng Zhang, Yong Chen, Jiayu. Gao, Bin Huang, Dengbo He, Shuiguang Deng, Hao Chen, Ying-Cong Chen
IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2026   (Best Paper Honorable Mention)
arXiv / bibtex

A driving copilot that adapts to each user's preferences and state.

Gene-M1: Advancing Cross-Species Genomic Discovery via Taxon-Specific Mixture-of-Experts
Yuhang Li*, Jiaqi Tang*, Jianmin Chen, Yourui Han, Xuequn Shang, Bolin Chen
International Symposium on Bioinformatics Research and Applications (ISBRA), 2026
bibtex

Taxon-specific experts for cross-species genomic discovery.

Responsive Test-Time Model Adaptation for Mobile Applications via Runtime-efficient Sparse Updates
Cheng Fang, Bin Guo, Sicong Liu, Zimu Zhou, Jiaqi Tang, Ke Ma, Shiyan Luo, Geyang Song, Zhiwen Yu
IEEE Transactions on Mobile Computing (TMC), 2026
bibtex

Sparse runtime updates keep mobile test-time adaptation responsive.

2025

RhythmGuassian: Repurposing Generalizable Gaussian Model For Remote Physiological Measurement
Hao Lu*, Yuting Zhang*, Jiaqi Tang, Bowen Fu, Wenhang Ge, Wei Wei, Kaishun Wu, Ying-Cong Chen
IEEE/CVF International Conference on Computer Vision (ICCV), 2025   (Highlight)
bibtex

A generalizable Gaussian model repurposed for contactless physiological sensing.

Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive Injection
Bowen Fu, Wei Wei, Jiaqi Tang, Jiangtao Nie, Xiaogang Xu, Ying-Cong Chen, Lei Zhang
IEEE/CVF International Conference on Computer Vision (ICCV), 2025
bibtex

Implicit style decoupling and adaptive injection for controllable stylization.

SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Sparsity Activation
Ke Ma, Jiaqi Tang, Fan Dang, Bin Guo†, Sicong Liu, Cheng Fang, Zhui Zhu, Lei Wu, Ying-Cong Chen, Zhiwen Yu, Yunhao Liu†
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025   (Highlight)
code / bibtex

Layer-wise activation pruning that cuts test-time adaptation memory.

Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
Qingfa Xiao, Jiachuan Wang, Haoyang Li, Cheng Deng, Jiaqi Tang, Shuangyin Li, Yongqi Zhang, Jun Wang, Lei Chen
arXiv preprint, 2025
arXiv / bibtex

Activation-guided probe queries retrieve KV pairs for long-context inference.

2024

AdaShadow: Responsive Test-time Adaptation for Non-stationary Mobile Environments
Cheng Fang, Sicong Liu, Zimu Zhou, Bin Guo†, Jiaqi Tang, Ke Ma, Zhiwen Yu
ACM Conference on Embedded Networked Sensor Systems (SenSys), 2024   (Best Paper Honorable Mention, Top 7/313)
bibtex

Selective layer updates give 2x to 3.5x faster on-device adaptation.

Hawk: Learning to Understand Open-World Video Anomalies
Jiaqi Tang*, Hao Lu*, Ruizheng Wu, Xiaogang Xu, Ke Ma, Cheng Fang, Bin Guo, Jiangbo Lu, Qifeng Chen, Ying-Cong Chen†
Annual Conference on Neural Information Processing Systems (NeurIPS), 2024
code / website / bibtex

Motion-aware video-language model for open-world anomaly understanding.

Learning to Remove Wrinkled Transparent Film with Polarized Prior
Jiaqi Tang, Ruizheng Wu, Xiaogang Xu, Sixing Hu, Ying-Cong Chen†
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
code / project page / bibtex

Polarization priors reconstruct surfaces occluded by wrinkled transparent film.

Scaling Multi-Camera 3D Object Detection through Weak-to-Strong Eliciting
Hao Lu, Jiaqi Tang, Xinli Xu, Xu Cao, Yunpeng Zhang, Guoqing Wang, Dalong Du, Hao Chen, Ying-Cong Chen†
arXiv, 2024
arXiv / code

Weak camera-specific experts elicit stronger multi-camera 3D detection.

An Incremental Unified Framework for Small Defect Inspection
Jiaqi Tang, Hao Lu, Xiaogang Xu, Ruizheng Wu, Sixing Hu, Tong Zhang, Tsz Wa Cheng, Ming Ge, Ying-Cong Chen†, Fugee Tsung
18th European Conference on Computer Vision (ECCV), 2024
code / project page / bibtex

Adds new products to a unified defect inspector without forgetting.

2023

High Dynamic Range Image Reconstruction via Deep Explicit Polynomial Curve Estimation
Jiaqi Tang, Xiaogang Xu, Sixing Hu, Ying-Cong Chen†
26th European Conference on Artificial Intelligence (ECAI), 2023   (Long Oral)
code / arXiv / talk / bibtex

Explicit polynomial tone-curve estimation for generalizable HDR reconstruction.

2021

NTIRE 2021 multi-modal aerial view object classification challenge
Jerrick Liu, Nathan Inkawhich, Oliver Nina, Radu Timofte, Sahil Jain, Bob Lee, Yuru Duan, Wei Wei, Lei Zhang, Songzheng Xu, Jiaqi Tang, and others
IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2021   (Winner Award, 1st Rank)
arXiv / bibtex

EO and SAR aerial object classification challenge, ranked first.

Surveys

2026

Intelligent Remote Sensing Agents: A Survey
Jiaqi Tang*, Yingying Yan*, Qianzhou Wang*, Yuyang Xia*, Botong Geng*, Jianmin Chen*, Ke Ma, Youyang Zhai, Qingfeng He, Weigeng Shao, Yunjin Sun, Junwei Dai, Chuxi Chen, Xiaogang Xu, Kelu Yao, Lei Zhang, Wei Wei†, Qifeng Chen†, Antonio Plaza, Yanning Zhang
Survey, 2026
paper / repository

Media Report: Synced (机器之心) / X / Reddit

A systematic review of 100+ works on intelligent remote sensing agents.

2024

GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
Hao Lu*, Xuesong Niu*, Jiyao Wang*, Yin Wang*, Qingyong Hu*, Jiaqi Tang*, Yuting Zhang, Kaishen Yuan, Bin Huang, Zitong Yu, Dengbo He, Shuiguang Deng, Hao Chen, Ying-Cong Chen†, Shiguang Shan
IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024   (Oral)
code / bibtex

Evaluating GPT-4V across five visual affective computing abilities.


← Back to Home