Large Language Models
166 views
Alignment Tax
Quick Definition
Performance cost of aligning AI with human preferences
Full Definition
The performance cost incurred when aligning AI models to follow human preferences and safety guidelines.
Examples
capability reduction after RLHF, safety vs helpfulness tradeoff
Related Terms
rlhf
ai-alignment
fine-tuning
More Large Language Models Terms
Function Calling
Structured mechanism for LLMs to generate API function calls
Synthetic Data Generation
Using AI to create artificial training data replicating patterns
QLoRA
Combining 4-bit quantization with LoRA for efficiency
Quantization LLM
Reducing LLM weight precision for efficient inference
Prompt Caching
Storing key-value states for repeated prompt prefixes
Structured Output
Generating LLM responses in predefined machine-readable formats