vlmrl
An experimental reinforcement learning gym for vision-language models, written in JAX. It lets you plug in environments and models and train them with PPO.
Share on XLicense: Apache-2.0
Overview
vlm-gym is an experimental reinforcement learning gym for vision-language models, written in JAX. The idea is to drop in any environment and any model and train with PPO. It ships with environments such as GeoGuessr, NLVR2 and captioning, and a reference Qwen3-VL-4B-Instruct model that can be converted from Hugging Face to JAX. The trainer, rollout engine and evaluation harness live in a core folder, and GeoGuessr training uses a staged curriculum.
Key features
- Pluggable vision environments: GeoGuessr, NLVR2, captioning
- PPO trainer for vision-language models
- Rollout engine and evaluation against a Hugging Face baseline
- Converts Qwen3-VL-4B-Instruct from Hugging Face to JAX
Best for
Suited to researchers experimenting with reinforcement learning on vision-language models in JAX. The README marks its status as experimental.
- Upstream
- sdan/vlm-gym
- Fork on GitHub
- Guo-astro/vlmrl
- Upstream stars
- 150
- Category
- AI agents and LLM tools
- Language
- Python
- License
- Apache-2.0
- Forked
- 2025-10-14
- Sync status
- In syncLast synced 2026-10-10
More in AI agents and LLM tools
An open-source curriculum for learning AI engineering by implementing model internals, retrieval pipelines, and agent runtimes. Study through lessons, interactive labs, or staged coding projects, and inspect the code and evaluation results.
Forked 2026-10-09Last synced 2026-10-10License: MITAI agents and LLM toolsGitHub
A small, hackable character-level language model that learns patterns from line-based text files and generates similar items, such as name ideas.
Forked 2026-10-09Last synced 2026-10-10License: MITAI agents and LLM toolsGitHub
uzu is an inference engine for running AI models in apps, helping developers keep inference on-device instead of relying on remote services. Use it to download a model and build chat features.
Forked 2026-10-09Last synced 2026-10-10License: MITAI agents and LLM toolsGitHub
Anti Slop gives AI coding agents rules to avoid generic UI, filler text, and AI-shaped code. Use it as a filter alongside your own design direction.
Forked 2026-10-08Last synced 2026-10-10License: MITAI agents and LLM toolsGitHub