{"description":"Test your understanding of BERT's architecture, masked language modeling, next sentence prediction, and how BERT processes and understands language.","questions":[{"answer":"Bidirectional Encoder Representations from Transformers","number":1,"options":["Bidirectional Encoder Representations from Transformers","Binary Entity Recognition Transformer","Basic Encoding and Recursive Translation","Back-End Recurrent Transformer"],"question":"What does BERT stand for?"},{"answer":"It reads text bidirectionally to capture full context","number":2,"options":["It only uses tokens from the first sentence","It processes words in random order","It reads text bidirectionally to capture full context","It uses convolutional layers instead of attention"],"question":"What makes BERT different from traditional left-to-right language models?"},{"answer":"Masked Language Modeling and Next Sentence Prediction","number":3,"options":["Sentiment prediction and summarization","Word embedding and decoding","Masked Language Modeling and Next Sentence Prediction","Grammar correction and translation"],"question":"What are the two pre-training tasks used by BERT?"},{"answer":"[CLS]","number":4,"options":["[SEP]","[PAD]","[CLS]","[MASK]"],"question":"Which token is used to classify the sentence pair in BERT?"},{"answer":"To mark a word that the model should predict","number":5,"options":["To separate two segments","To identify padding locations","To mark a word that the model should predict","To initialize a new sentence"],"question":"What is the purpose of the [MASK] token during training?"},{"answer":"It helps the model focus on different parts of the input simultaneously","number":6,"options":["It encodes token frequency","It helps the model focus on different parts of the input simultaneously","It replaces position embeddings","It outputs class probabilities"],"question":"In BERT\u2019s architecture, what role does the Multi-Head Attention mechanism play?"},{"answer":"A flag indicating if one sentence follows another","number":7,"options":["A list of part-of-speech tags","A flag indicating if one sentence follows another","An integer index of masked tokens","A count of sentence pairs"],"question":"What is the output of the 'IsNext' classification task in BERT?"},{"answer":"They stabilize and encode richer semantic meaning","number":8,"options":["They become identical across all tokens","They stabilize and encode richer semantic meaning","They decrease in dimensionality","They are discarded after each epoch"],"question":"What happens to embedding vectors after many training epochs?"},{"answer":"To stabilize learning and avoid vanishing gradients","number":9,"options":["To speed up translation","To stabilize learning and avoid vanishing gradients","To align token indices","To reduce vocabulary size"],"question":"Why are residual connections and layer normalization used in BERT?"},{"answer":"By passing encoded vectors through a decoder and applying softmax","number":10,"options":["Using a lookup table from the first encoder","By matching frequency distributions","By passing encoded vectors through a decoder and applying softmax","By averaging all tokens"],"question":"How does BERT generate final predictions for masked tokens?"}],"title":"BERT Explorer"}
