CS @ UC Berkeley | LLM Research | Full-Stack
I build things at the intersection of machine learning and real products.
FAME Benchmark — co-authored an LLM evaluation benchmark accepted to the ICLR 2025 Workshop Logical Reasoning of Large Language Models, in collaboration with researchers from Algoverse. Evaluated GPT-4.1, Gemini 2.5 Pro, Claude 3.5 Sonnet, LLaMA 4, and Qwen 3 on logical reasoning under false mathematical facts.
- Film Compass — ML-powered movie discovery app. 5,000+ films embedded, clustered, and visualized as an interactive scatter plot using GPT-4o-mini,
mxbai-embed-large-v1, Leiden clustering, and UMAP. - Cuisync — iOS app with 200+ downloads. Voice-driven recipe assistant built with Faster Whisper and FastAPI.
- FAME — LLM benchmark for evaluating reasoning under false mathematical facts.
Python · PyTorch · Hugging Face · FastAPI · TypeScript · React Native · C++ · Supabase