GTC‑2.5V Preview extends the GTC series with multimodal capabilities. Integrating a lightweight visual encoder with a 64M‑parameter decoder‑only backbone, the model supports image captioning, visual question answering, and text reading from images. Trained end‑to‑end on a V100 32G GPU, it validates TAI Research's multimodal pipeline and serves as a baseline for larger vision‑language models.

As a research preview, it is intended for technical validation and internal benchmarking rather than production use.

We thank the MiniMind‑V team for their foundational open‑source work.

📄 Download System Card (PDF)