vLLM Project
Jul 2026 - Present·Open Source
Open Source Contributor
Contributed to 20+ merged PRs in the vLLM Metal backend, spanning speculative decoding, paged KV-cache management, model compatibility, sampling, and memory management for Apple Silicon inference.
Implemented DSpark speculative decoding for the vLLM Metal backend, including Qwen3 model integration, scheduler-owned KV-cache management, paged execution, Markov drafting, prefix caching, validation, and serving benchmarks.
Optimized speculative decoding by eliminating redundant KV re-ingestion and chunking cold draft-KV ingestion: 31× faster 8K first-propose latency (3.47s → 112ms), 2.1× end-to-end latency (64.0 → 30.8 ms/token), 50% less steady-state KV ingestion, and 48% lower DSpark draft-weight memory (4.42 → 2.29 GiB).
Amazon
Jan 2026 - Jul 2026·Berlin, Germany
Software Development Engineer Intern
Engineered a full-stack embedding search and visualization platform over 50M embeddings, with PCA/UMAP visualization and sub-5ms K-Nearest-Neighbours queries, accelerating Music recommender experimentation.
Designed a batch ML data pipeline for embedding metadata hydration and model indexing across S3, ECS, and OpenSearch, enabling automated weekly model updates.
Built evaluation workflows measuring marketplace performance with Recall, NDCG, Precision, and triplet accuracy, alongside a FastAPI backend and CloudWatch monitoring.
KBDfans (Partner)
Mar 2024 - Dec 2025·Changzhou, Jiangsu, China
Firmware Lead & Embedded Software Engineer
Led firmware and hardware development for the Zellia Hall Effect project (team of five C/C++ engineers): a distributed embedded system across five AT32 MCUs with enhanced modularity and multi-layout keyboard support, designed in KiCad 8.0.
Achieved <0.28 ms input latency at a 106 kHz scan rate by decoupling signal acquisition and processing, with a custom 7.5 Mbps UART protocol for high-throughput multi-MCU synchronization.
Offloaded ADC normalization and calibration to slave MCUs, reducing master CPU load by 90%.