Attention Score
Computed similarity between query and key vectors before softmax normalization in attention.
In italiano: Attention ScoreScore = dot_product(Query, Key) / sqrt(d_k). Scaled dot-product prevents gradient issues. Higher scores mean more relevance.
Examples
- Scaled dot-product attention
- Query-key similarity
- Attention computation