Artificial Intelligence
24 views
Reward Modeling
Quick Definition
Training models to predict human preferences for RLHF
Full Definition
Training a model to predict human preferences and provide reward signals for reinforcement learning from human feedback.
Examples
RLHF training pipeline, preference learning, alignment
Related Terms
rlhf
ai-alignment
reinforcement-learning
More Artificial Intelligence Terms
Adversarial Robustness
AI model ability to resist adversarial perturbations
Hallucination
AI generating confident but factually incorrect information
In-Context Learning
LLMs learning new tasks from prompt examples without weight updates
Neural Architecture Search
Automating the design of optimal neural network architectures
Meta-Learning
Designing models that learn how to learn new tasks
Grounding
Connecting AI outputs to verifiable real-world sources