You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Browser LLM inference on hand-written WGSL — up to a 35B sparse MoE, and models too big for one machine split by layer range across several over WebRTC. No TVM, no WebLLM runtime. Every number in BENCH.md.
High-performance, zero-dependency JS/TS Language Model & NLP suite for Node.js and browsers featuring text generation, self-attention heatmaps, RAG fact injection, and webpage autocomplete.
Scaffold for browser-local LLM chat apps. Runs Gemma 4 entirely in the browser via MediaPipe + WebGPU — no server, no cloud, no tokens. Mobile-first PWA, Atomic Design, i18n (DE/EN), multi-conversation. Fork it. Own it. Ship it.
A zero-dependency Node.js HTTPS-to-HTTP proxy that lets browser-based AI tools (Agile V Studio, Framewrk Studio) talk to local LLM backends like Ollama without mixed-content errors. Includes self-signed TLS, CORS, token tracking dashboard, Docker support, and OpenAI/Anthropic-compatible passthrough.
Terminal-style portfolio website that runs a local LLM in your browser. Type `ask` to chat with an on-device LLM (Gemini Nano / WebLLM via WebGPU) — no backend, no API keys, your messages stay local. Built with React + Vite.