I'm passionate about AI Infrastructure and Open Source, especially the journey from GPU → LLM Inference → Tokens.
I contribute to projects across the AI infrastructure stack, including vLLM, SGLang, llm-d, KServe, HAMi, and Dynamo, with a focus on Kubernetes, GPU infrastructure, inference runtimes, scheduling, routing, and LLM serving.
I also share what I learn through my 📖 blog. Always happy to connect, collaborate, and learn from the community! 🚀
Currently learning Spanish on Duolingo. ¡Hola! 🇪🇸





