Large Language Models
20 views
Reward Model
Quick Definition
Model scoring LLM outputs based on human preferences
Full Definition
A model trained to score LLM outputs based on human preference data for RLHF training pipelines.
Examples
RLHF pipeline, preference ranking, reward shaping
Related Terms
rlhf
reward-modeling
ai-alignment
More Large Language Models Terms
Fine-Tuning
Further training pre-trained models on specific domain data
Model Merging
Combining fine-tuned LLMs without additional training
RAG
Enhancing LLMs by retrieving relevant external knowledge
QLoRA
Combining 4-bit quantization with LoRA for efficiency
Structured Output
Generating LLM responses in predefined machine-readable formats
Guardrails LLM
Frameworks for monitoring and controlling LLM I/O