Large Language Models
32 views
Alignment Tax
Quick Definition
Performance cost of aligning AI with human preferences
Full Definition
The performance cost incurred when aligning AI models to follow human preferences and safety guidelines.
Examples
capability reduction after RLHF, safety vs helpfulness tradeoff
Related Terms
rlhf
ai-alignment
fine-tuning
More Large Language Models Terms
Speculative Decoding
Using draft model candidates verified by large model
RLHF
Aligning LLMs with human preferences using reinforcement learning
Scaling Laws
Relationships between model size, data, compute, and performance
Beam Search
Search algorithm exploring multiple output sequences
Structured Output
Generating LLM responses in predefined machine-readable formats
Next Token Prediction
Core LLM objective predicting the next token from context