Artificial Intelligence
141 views
Reward Modeling
Quick Definition
Training models to predict human preferences for RLHF
Full Definition
Training a model to predict human preferences and provide reward signals for reinforcement learning from human feedback.
Examples
RLHF training pipeline, preference learning, alignment
Related Terms
rlhf
ai-alignment
reinforcement-learning
More Artificial Intelligence Terms
Quantization
Reducing model precision from 32-bit to lower bit representations
AI Safety
Research ensuring AI is developed safely without harm
Deep Learning
Machine learning using multi-layered neural networks
Sentiment Analysis
NLP technique determining emotional tone in text
Zero-Shot Learning
Performing tasks on unseen classes without training examples
One-Shot Learning
Learning to recognize patterns from a single example