AI Infrastructure Engineer | LLM Researcher | Embedded Systems Specialist
| Stack | Technologies |
|---|---|
| LLM | llama.cpp, GGUF, Qwen3.5, LoRA, MoE routing, low-VRAM optimization |
| Infrastructure | NVIDIA Tesla P40 (sm_61), RTX 3050, CUDA workarounds, Docker, Tailscale |
| Embedded | ESP32, Arduino R4 WiFi, MQTT, HID, OpenClaw gateway |
| Automation | Bash pipelines, systemd, Playwright, Obsidian vault hooks |
- mini-phase-twin-30b-low-vram-gguf – Low-VRAM GGUF model for consumer GPUs (Tesla P40)
- auto-quantization-pipeline-gguf – Automated benchmarking and quantization for Q4_K_M/Q5_K_S
- 4-agent-wrappers-on-qwen3.6-27b – Multi-agent routing and inference optimization
- ai-gateway-in-prod-alternative-concrete-a-litellm – OpenAI-compatible gateway for NVIDIA consumer hardware
- ai-home-assistant-hid-dashboard – Physical dashboard for local AI stack monitoring
- ai-model-selector-physical-controller – ESP32-based model selection interface
- auto-vault-journal – Obsidian vault auto-update via Claude Code hooks
- dictate – Local Whisper dettatura for Claude Code (Italian support)
- add-video-input-support-to-llamacpp-mtmd – Video frame acquisition for LLM inference
- ai-influencer-pubblicazione-social – Social media pipeline with LoRA face-consistency
- ai-dashboard – Web dashboard for GPU monitoring and task automation