Transformer
Quick Definition
Architecture using self-attention for parallel processing
Full Definition
A neural network architecture using self-attention mechanisms for parallel processing of sequential data.
Examples
GPT, BERT, Vision Transformer
Related Terms
self-attention
large-language-model
More Artificial Intelligence Terms
AutoML
Automation of the machine learning pipeline
Few-Shot Learning
Learning to perform tasks from a small number of examples
State Space Model
Architecture using linear attention for efficient sequence modeling
Convolutional Neural Network
Neural network for image and grid-like data processing
Pruning
Removing redundant neural network connections to reduce size
Retrieval-Augmented Generation
Combining LLMs with external knowledge for accurate responses