Inference-time scaling on Red Hat AI: Improving model reliability
How inference-time scaling techniques on Red Hat AI improve model reliability by letting models spend more compute at generation time on harder problems.
A collective of researchers and engineers from Red Hat & IBM building LLM toolkits you can use today.
Inference-time scaling for LLMs.
Synthetic data generation pipelines
Post training algorithms for LLMs
Asynchronous GRPO for scalable reinforcement learning.
A method for skipping redundant attention blocks in language models
Efficient training library for large language models up to 70B parameters on a single node.
Adaptive SVD-based continual learning method for LLMs.
Inference-time scaling with particle filtering.
State-of-the-art reward models for preference data generation and acceptance criteria.
KV cache quantization for scaling inference time
Efficient messages-format SFT library for language models
How inference-time scaling techniques on Red Hat AI improve model reliability by letting models spend more compute at generation time on harder problems.
Particle filtering offers a principled alternative to majority voting and beam search for inference-time scaling, improving LLM reasoning by pruning unpromising chains of thought early.
Orthogonal Subspace Fine-Tuning (OSFT) prevents catastrophic forgetting during LLM fine-tuning by constraining weight updates to directions that preserve existing model capabilities.
📹 Scaling Training and Eval Data for Coding Agents
👤 Speaker: Ari Aye
August 07, 2026
📹 Policy-Driven Agentic Red-teaming
👤 Speaker: Muneeza Azmat
July 31, 2026
📹 Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
👤 Speaker: Rachel Ma
July 17, 2026