{"description":"Test your understanding of attention heatmaps, multi-head attention visualization, and how attention scores evolve during model training.","questions":[{"answer":"Relationships and focus between words based on attention scores","number":1,"options":["Token frequency","Word embeddings","Relationships and focus between words based on attention scores","Model accuracy over time"],"question":"What does an attention heatmap visualize in NLP models?"},{"answer":"Words attending and words being attended to","number":2,"options":["Document length and punctuation","Input and output token embeddings","Words attending and words being attended to","Layer count and neuron ID"],"question":"What do the rows and columns represent in an attention heatmap?"},{"answer":"A higher attention score between two tokens","number":3,"options":["More padding in the sequence","A higher attention score between two tokens","The presence of a special token","A longer input sentence"],"question":"What does a brighter color in an attention heatmap cell typically indicate?"},{"answer":"To capture different relationships and features across multiple perspectives","number":4,"options":["To speed up tokenization","To reduce the embedding size","To capture different relationships and features across multiple perspectives","To improve dropout rate"],"question":"Why do Transformer models use multi-head attention?"},{"answer":"Each head focuses on a different aspect of word relationships","number":5,"options":["Each head focuses on a different aspect of word relationships","They determine the sentence length","They mark punctuation importance","They replace position embeddings"],"question":"What is the role of attention heads in a heatmap visualization?"},{"answer":"It shows sharper and more meaningful focus areas","number":6,"options":["It becomes more random","It shows sharper and more meaningful focus areas","It fades to zero values","It removes less frequent words"],"question":"How does training affect the attention heatmap over time?"},{"answer":"Dot products between queries and keys","number":7,"options":["Convolution operations","RNN cell outputs","Dot products between queries and keys","One-hot vector comparisons"],"question":"What is typically used to compute attention scores in BERT?"},{"answer":"By concatenation followed by a linear transformation","number":8,"options":["By averaging","By summing","By concatenation followed by a linear transformation","By multiplying with position embeddings"],"question":"How are attention head outputs typically combined in a Transformer?"},{"answer":"Masked or content-heavy tokens","number":9,"options":["Only [PAD] tokens","Stop words only","Masked or content-heavy tokens","Output tokens"],"question":"Which tokens in a sentence are often used to evaluate attention patterns?"},{"answer":"It visualizes how models distribute focus across input tokens","number":10,"options":["It predicts masked tokens","It visualizes how models distribute focus across input tokens","It measures GPU usage","It increases training speed"],"question":"Why is the attention heatmap a useful interpretability tool in NLP?"}],"title":"Attention Heatmap"}
