RLHF
Quick Definition
Aligning LLMs with human preferences using reinforcement learning
Full Definition
Reinforcement Learning from Human Feedback aligning language models with human preferences.
Examples
ChatGPT alignment, preference learning, safety tuning
Related Terms
reinforcement-learning
reward-model
alignment
More Large Language Models Terms
Grounded Generation
Generating outputs faithful to provided source documents
Function Calling
Structured mechanism for LLMs to generate API function calls
Constitutional AI
Training AI to follow explicit principles for consistency
Prompt Template
Reusable structured format with dynamic input placeholders
GPT
OpenAI's family of autoregressive large language models
Embeddings LLM
Dense vectors capturing semantic meaning of text