📚 Auto-publish: Add/update 4 blog posts
Some checks failed
Hugo Publish CI / build-and-deploy (push) Has been cancelled
Some checks failed
Hugo Publish CI / build-and-deploy (push) Has been cancelled
Generated on: Sat Aug 16 21:13:18 UTC 2025 Source: md-personal repository
This commit is contained in:
@@ -9,7 +9,7 @@ Large Language Models (LLMs) have demonstrated astonishing capabilities, but out
|
||||
|
||||
You may have seen diagrams like the one below, which outlines the RLHF training process. It can look intimidating, with a web of interconnected models, losses, and data flows.
|
||||
|
||||
![[Pasted image 20250816140700.png]]
|
||||

|
||||
|
||||
This post will decode that diagram, piece by piece. We'll explore the "why" behind each component, moving from high-level concepts to the deep technical reasoning that makes this process work.
|
||||
|
||||
|
Reference in New Issue
Block a user