Shaobo Guo (郭少博). I'm currently working as a full-time algorithm engineer at Bytedance.
My work focuses on areas such as computer vision and multi-modal large language models. I have extensive experience in applying advanced algorithms to real-world business scenarios, building end-to-end systems that directly support and optimize company operations. Over the past few years, I have led or participated in multiple projects that significantly improved efficiency and solved complex business problems through intelligent algorithmic design and system-level implementation.
Conference Publications
ICCV 2025 — Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Tech Reports
Mammothmoda1.0 — Mammothmoda: Multi-modal large language model
Mammothmoda2.0 — Mammoth2: A Unified AR Diffusion Framework for Multimodal Understanding and Generation