After 220+ hours of training across 24 sessions, HTGM v2 Hindi LLM is now approaching the end of its base pre-training phase.
This is not an official final stop yet, but training may be paused to avoid overfitting. If needed, training can still continue further.
- ~9GB Hindi dataset
- Experiment phase
- Learning training pipeline
- ~41GB Hindi dataset
- Foundation phase
- 220+ hours training
- ~163M parameter GPT-style model
Now the next major step is:
The upcoming work will focus on:
- QA datasets
- Instruction-following
- Better response formatting
- Improved reasoning
- Higher answer accuracy
HTGM v3 is currently a future research vision.
Target:
- 1TB → 10TB high-quality Hindi dataset
This would be a massive scale jump compared to:
- HTGM v1 → 9GB
- HTGM v2 → 41GB
Kaggle GPUs were enough for experimentation and learning.
But future HTGM v3 training may require:
- NVIDIA A100 GPUs
- NVIDIA H100 GPUs
- Enterprise-level infrastructure
- Multi-GPU systems
Support from startups, companies, or research organizations may be needed in the future.
HTGM v3 is NOT starting now.
There is still a lot of work left:
- Better datasets
- Better optimization
- Better fine-tuning
- Better architecture
Right now, the focus remains on improving HTGM v2.
HTGM v1 → Experiment
HTGM v2 → Foundation
HTGM v3 → Scale & Intelligence
Small steps today. Bigger AI tomorrow.
— Mahesh Editor