Attention Score

Computed similarity between query and key vectors before softmax normalization in attention.

In italiano: Attention Score

Score = dot_product(Query, Key) / sqrt(d_k). Scaled dot-product prevents gradient issues. Higher scores mean more relevance.

Examples

  • Scaled dot-product attention
  • Query-key similarity
  • Attention computation