Improvement Technique

Improvement Technique

LFM2.5-2.6B: Deploy Agents Everywhere — Blog
LFM2.5-2.6B: Deploy Agents Everywhere — Blog
LFM2.5-2.6B is an on-device agentic model that plans, calls tools, and runs multi-step tasks at 220 tok/s in under 2.5 GB. Open weights on Hugging Face.
MOPD. We then use the specialized experts as teachers and distill their capabilities into a single student model. Unlike off-policy distillation, where the student learns from trajectories generated by another model, MOPD lets the student roll out under its own policy. Each prompt is routed to the teacher for its corresponding domain, which supervises the student's response with token-level feedback.
·liquid.ai·
LFM2.5-2.6B: Deploy Agents Everywhere — Blog
Fine-tuning Small Language Models
Fine-tuning Small Language Models
How a fine-tuned 3-billion-parameter model beats today’s frontier models at one job, and how to build one this weekend
·open.substack.com·
Fine-tuning Small Language Models
Quantization · Hugging Face
Quantization · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Quantization
·huggingface.co·
Quantization · Hugging Face
We’re introducing HALO 😇
We’re introducing HALO 😇
Hierarchal Agent Loop Optimizer HALO is an RLM-based agent optimization technique capable of recursively self-improving agents by analyzing their execution traces and suggesting changes. This work is inspired by the Mismanaged Genius Hypothesis
·x.com·
We’re introducing HALO 😇
Understanding RAG vs Fine-Tuning
Understanding RAG vs Fine-Tuning
Discover the key differences between RAG and fine-tuning, what each approach can bring, and how to choose the right AI approach for your business goals.
·cohere.com·
Understanding RAG vs Fine-Tuning
LangGraph Rollout: Evolving VeRL’s Multi-Turn Capabilities for Agent RL
LangGraph Rollout: Evolving VeRL’s Multi-Turn Capabilities for Agent RL
After completing our multi-turn tokenization and masking refactoring, we eliminated a critical bottleneck that was preventing us from building a more consistent and flexible rollout system for our Agent RL research. This breakthrough enabled us to implement a LangGraph-based rollout for VeRL in just a few days, which we’ve already successfully deployed in our Agent RL experiments. In this article, I’ll share our journey from VeRL’s native multi-turn implementation to our new LangGraph-based solution, explaining both the motivations driving this evolution and the technical details of our implementation.
·jybsuper.github.io·
LangGraph Rollout: Evolving VeRL’s Multi-Turn Capabilities for Agent RL
"regular people don't fine-tune VLMs"
"regular people don't fine-tune VLMs"
but wtf not? - skill gap - high fine-tuning costs - lack of standards and unified approaches over the past few weeks I've been working on maestro - streamlined tool for VLM fine-tuning link: — SkalskiP (@skalskip92)
·x.com·
"regular people don't fine-tune VLMs"