Synth AI
WorkshopCookbooksBlogDocsStackStackSign in
WorkshopCookbooksBlogDocsStackStackSign in
Changelog

Terminal Training Logs & Qwen-VL RL

Friday product updates for October 31, 2025.

Friday, October 31, 20251 min read
traininglogsverifiersqwen-vl
Archive
This is historical material. Use the current Managed Research, Research Factory, and GEPA/GELO pages for product decisions.

TL;DR

  • Terminal Training Logs: Full real-time streaming logs for SFT and RL training
  • Hosted Verifiers: Configurable Synth verifiers with per-job overrides
  • Qwen-VL Support: Vision models now supported across SFT & RL
  • Rubric-Aware Filtering: SFT filtering pipelines with structured rubric definitions

Terminal Training Logs

Both uvx synth-ai train for SFT and RL now provide comprehensive real-time training logs directly in the terminal.

Features

  • Live Status Updates: See QUEUED, RUNNING, and other status updates in real-time
  • Detailed Event Logs: Timestamps and sequence numbers for all events
  • Full Metrics Logging: Training loss, learning rate, GPU utilization, KL divergence, rollout times
  • Timeline Progression: Visual timeline showing progress throughout the entire training process

Rubrics, Hosted Verifiers & Qwen-VL RL

Hosted Synth Verifiers

Rollout filtering and on-policy RL can now invoke hosted verifiers with per-job overrides:

  • Rubric Selection: Choose from Synth-hosted rubrics for consistent evaluation
  • Concurrency Caps: Control how many verifier evaluations run concurrently
  • Strict Behavior: Configure strict behavior when verifiers are unavailable

Rubric-Aware Filtering

SFT filtering pipelines accept structured rubric definitions:

  • Structured Scoring: Traces are scored according to your rubric criteria
  • Automatic Trimming: Traces are trimmed before export based on rubric scores
  • Custom Criteria: Define your own evaluation criteria for filtering

Qwen-VL Support

Qwen3-VL models can be fine-tuned and trained with RL:

  • Vision Collators: Built-in vision collators for image processing
  • LoRA Projector Targeting: LoRA adapters target vision projectors
  • Rollout Plumbing: Full support for vision models in RL rollouts

Instruct-Model RL Guidance

Added documentation and defaults for running RL on Qwen instruct SKUs:

  • Semaphore Tuning: Avoid premature episode completion
  • Best Practices: Guidance on configuring RL for instruct models

Documentation

  • RL Documentation: Updated RL guides with Qwen-VL examples
  • Verifier Configuration: Documentation for configuring hosted verifiers
  • Rubric Guide: Guide for creating and using rubrics in filtering pipelines

Use Cases

  • Real-Time Monitoring: Monitor training progress directly in terminal without switching contexts
  • Quality Filtering: Use rubric-based filtering to improve training data quality
  • Vision RL: Train RL models on vision tasks with Qwen-VL
  • Consistent Evaluation: Use hosted verifiers for consistent evaluation across experiments
←Back to Changelog

Ready to try Synth?

Get started with serverless RL training and prompt optimization.

Get StartedSchedule Demo
On this page
  • TL;DR
  • Terminal Training Logs
  • Features
  • Rubrics, Hosted Verifiers & Qwen-VL RL
  • Hosted Synth Verifiers
  • Rubric-Aware Filtering
  • Qwen-VL Support
  • Instruct-Model RL Guidance
  • Documentation
  • Use Cases
© 2026 SynthWorkshopReleasesCookbooksChangelogOpen sourceDocsBook a Demo