I build and ship predictive models — for predictive maintenance and prognostics, where calibration and reliability under drift matter more than benchmark scores.
The work I'm doing on my own time is on LLM agents and evaluation: what makes a small model reliable, how you measure it, and how much of agent performance comes from the harness rather than the model. That's what hiveloom is — generate, run, and evolve agent harnesses so small, cheap models do repeatable, verifiable work.
Happy to talk to anyone working or interested in similar problems!
Links
- LinkedIn: in/francesco-mrn
- X: @francesco_mrno



