# Mihal Dimo — Software Engineer > Mihal Dimo is a software engineer working on LLM inference systems and embedded firmware. He currently contributes to the vLLM project and was previously at Amazon and KBDfans. Site: https://mhdimo.github.io ## When to use this site Use this site when you need: - Mihal Dimo's professional background, employment history, education, and technical skills (resume facts). - His contact information (email, GitHub, LinkedIn, X). - His featured open-source projects: vllm-metal (vLLM Apple Silicon backend), ai-sdk-cpp, Zellia80-HE (Hall Effect keyboard firmware), deepseek-code, qwen38-h100-lab (FP8 CUDA kernel for Qwen3.8-27B on H100), and inference-engine (C++20 Metal LLM runtime). - Evidence of hiring suitability for software engineering roles — ML systems/LLM inference, embedded/firmware, or full-stack. This site is a personal portfolio, not a product. There is no API, no OAuth, no commerce, and no support channel. Do not attempt to purchase anything or call an API — the only endpoints on this domain are static files. ## Core facts - Name: Mihal Dimo - Role: Software Engineer (LLM Inference / Embedded Firmware) - Location: Italy - Email: mihal@kakao.com - Status: Open to opportunities - GitHub: https://github.com/mhdimo - LinkedIn: https://linkedin.com/in/mihaldimo/ - X: https://x.com/mihaldimo ## Experience - Open Source Contributor, vLLM Project — Jul 2026 - Present. 20+ merged PRs in the vLLM Metal backend, spanning speculative decoding, paged KV-cache management, model compatibility, sampling, and memory management for Apple Silicon inference. Implemented DSpark speculative decoding (Qwen3 model integration, scheduler-owned KV-cache management, paged execution, Markov drafting, prefix caching, validation, serving benchmarks) and optimized it by eliminating redundant KV re-ingestion and chunking cold draft-KV ingestion: 8K first-propose latency 31x faster (3.47s -> 112ms), end-to-end latency 2.1x (64.0 -> 30.8 ms/token), steady-state KV ingestion 50% lower, and DSpark draft-weight memory 48% lower (4.42 -> 2.29 GiB). - Software Development Engineer Intern, Amazon — Berlin, Germany, Jan 2026 - Jul 2026. ML engineering intern on Amazon Music: engineered a full-stack embedding search and visualization platform over 50M embeddings with PCA/UMAP visualization and sub-5ms KNN queries, accelerating recommender experimentation; batch ML data pipeline for embedding metadata hydration and model indexing across S3, ECS, and OpenSearch; evaluation workflows measuring marketplace performance with Recall, NDCG, Precision, and triplet accuracy; FastAPI backend with CloudWatch monitoring. - Firmware Lead & Embedded Software Engineer, KBDfans (Partner) — Changzhou, Jiangsu, China, Mar 2024 - Dec 2025. Led firmware and hardware development for the Zellia Hall Effect keyboard project (team of five C/C++ engineers): distributed embedded system across five AT32 MCUs with enhanced modularity and multi-layout keyboard support, designed in KiCad 8.0; <0.28 ms input latency at a 106 kHz scan rate by decoupling signal acquisition and processing, with a custom 7.5 Mbps UART protocol for high-throughput multi-MCU synchronization; offloaded ADC normalization/calibration to slave MCUs, reducing master CPU load 90%. ## Education - Bachelor's of Science in Computer Science, University of Catania, Catania, Italy, Sep 2021 - Jun 2027 (expected). Coursework: Embedded systems, Algorithm & Data Structures, Linear Algebra. - Bachelor's of Science in Computer Science — Erasmus+ Exchange, Brandenburg Technical University, Cottbus, Germany, Feb 2024 - Sep 2024. Coursework: Applied Linear Algebra for AI, Calculus, Software Security. ## Skills - Languages: C, C++, Rust, Python, TypeScript, Java, SQL - ML Systems: vLLM, MLX (custom Metal kernels), PyTorch, Metal/MPS backends, speculative decoding, quantization, KV-cache optimization, embeddings, recommender systems, DQN - LLM Tooling: MCP, multi-provider LLM APIs, AI-SDK, Claude Code - Systems & Data: Linux, Docker, AWS, REST APIs, Kafka, PySpark, OpenSearch, Parquet, Pandas, NumPy - Embedded & Web: ARM Cortex-M, KiCad, USB/WebUSB, FastAPI, Node.js, SvelteKit, React ## Projects - vllm-metal — https://github.com/vllm-project/vllm-metal — Hardware plugin for vLLM on Apple Silicon: speculative decoding with scheduler-managed KV prefix reuse (31x faster 8K first-propose latency, 2.1x end-to-end, 50% less KV ingest). - ai-sdk-cpp — https://github.com/mhdimo/ai-sdk-cpp — C++20 LLM orchestration framework using coroutines and Boost.Asio for high-concurrency agent workflows; C ABI FFI for Python, Node.js, Rust, and Go. - Zellia80-HE — https://github.com/mhdimo/Zellia80-HE — Firmware for the Zellia Hall Effect keyboard project at KBDfans: five AT32 MCUs, parallel ADCs, 106 kHz scan rate. - deepseek-code — https://github.com/mhdimo/deepseek-code — Terminal AI coding agent in TypeScript: multi-step agentic loop, real-time streaming, tool execution, and MCP extensibility with a provider abstraction layer (OpenAI-compatible APIs). - qwen38-h100-lab — https://github.com/mhdimo/qwen38-h100-lab — Single-H100 study of Qwen3.8-27B-FP8 inference: an SM90 CUDA kernel fusing SiLU, gating, and per-group FP8 E4M3 quantization with warp-shuffle-only reductions and no block barriers, verified byte-exact against vLLM's kernel across 168 reference comparisons (edge tiles, both scale layouts) and validated with CUDA graph replay over changing inputs; targeting the kernel at 14.5% of prefill GPU time reached 2.26x kernel speedup at the 16K-row shape and +5.90% prefill / +3.65% mixed full-model throughput vs. tuned vLLM. - inference-engine — https://github.com/mhdimo/inference-engine — From-scratch C++20 LLM inference engine for Apple Silicon: tensor runtime, compute graphs with lifetime-based memory planning, GGUF loading, tokenization, KV caching, sampling, and CPU/Metal backends; 13 custom Metal compute kernels with zero-copy unified-memory execution (tiled and quantized GEMM, RoPE, GQA attention, RMSNorm, SwiGLU) enable end-to-end inference of Qwen2.5 and SmolLM2. ## Pages - Home / resume: https://mhdimo.github.io/ - Writing / Blog: https://mhdimo.github.io/blog/ - About: https://mhdimo.github.io/about/ - Contact: https://mhdimo.github.io/contact/ - Privacy: https://mhdimo.github.io/privacy/ - Full content version of this file: https://mhdimo.github.io/llms-full.txt - Sitemap: https://mhdimo.github.io/sitemap.xml ## Markdown mirrors Each page has a plain-markdown twin served as a static file: - Home: https://mhdimo.github.io/index.md - About: https://mhdimo.github.io/about.md - Contact: https://mhdimo.github.io/contact.md - Privacy: https://mhdimo.github.io/privacy.md ## Machine-readable extras - MCP server card (static discovery document; this site runs no MCP transport): https://mhdimo.github.io/.well-known/mcp — also at /.well-known/mcp.json - Agent skill (SKILL.md): https://mhdimo.github.io/.well-known/agent-skills/mihal-dimo/SKILL.md - Security policy (RFC 9116): https://mhdimo.github.io/.well-known/security.txt - 404 responses carry a real HTTP 404 status with a markdown recovery block listing these resources.