1. Quel est le rôle principal du Driver Program dans l’architecture distribuée de Spark ?
2. Quel composant convertit une colonne catégorielle en indices numériques appris lors du fit ?
3. Quel est le schéma de déploiement décrit pour l’inférence en deep learning avec Spark et Keras ?
Big Data — caractéristiques principales ?
Volume, Vélocité, Variété, Véracité, Valeur
Écosystème Spark — composantes clés ?
Spark SQL, Spark Streaming, MLlib, GraphX
Architecture Spark — modèle ?
Master/Slave avec Driver, Cluster Manager, Executors
Lazy Evaluation — mécanisme ?
Transformations enregistrées, actions déclenchent l’exécution
RDD — définition ?
Resilient Distributed Dataset, immuable et distribué
Transformations — rôle ?
Créer un nouveau RDD sans exécution immédiate
The revision sheet covers the essential concepts of Introduction à Spark et Big Data. It is organized by topic to facilitate learning and memorization, with key definitions, explanations and summaries.
Read the full sheet →The quiz contains 20 multiple-choice questions with detailed corrections and explanations for each answer. Ideal for testing your knowledge and identifying gaps.
Take the quiz (20 questions) →Revizly offers 20 interactive flashcards on Introduction à Spark et Big Data. Each card presents a question on the front and the answer on the back, enabling active and effective revision based on spaced repetition.
See all 20 flashcards →Import your PDF or paste your course, AI generates sheets, quizzes and flashcards in 30 seconds.