Microsoft 70-775 - Perform Data Engineering on Microsoft Azure HDInsight Exam

Question #11 (Topic: )
and three
data nodes. You have a MapReduce job.
You receive a notification that a data node failed.
to identity which component caused the failure.
Which tool should you use?
A. Job Tracker B. TaskTracker C. ResourceManager D. ApplicationMaster
Answer: C
Question #12 (Topic: )
Note: This question is part of a series of questions that present the same Scenario.
Each question I the series contains a unique solution that might meet the stated
goals. Some question sets might have more than one correct solution while others
might not have correct solution.
Start of Repeated Scenario:
initial data that contains the crime data from major cities.
You plan to build training models from the training data. You plan to automate the process
of adding more data to the training models and to training the models by using the
additional data, including data that is collected in near real time. The system will be used to
analyze event data gathered from many different sources. Such as Internet of things (IoT)
increased crime risk at a particular time and ptace.
incoming data stream from
Facebook. which are event-based only, rather than time-based. You also have a time
interval stream every 10 seconds.
number that defines
how many times a hashtag occurs within a Facebook post or how many times a tweet that
contains a specific hashtag is retweeted.
You must use the appropriate data storage, stream analytics techniques, and Azure
HDInsight cluster types tor the various tasks associated to the processing pipeline.
End of repeated Scenario.
You are designing the real-time portion of the input stream processing. The input will be a
continuous stream of data and each record will be processed one at a time. The data will
come from an Apache Kafka producer.
You need to identify which HDInsight cluster to use for the final processing of the input
data. This will be used to generate continuous statistics and real-time analytics. The
latency to process each record must be less than one millisecond and tasks must be
performed in parallel.
Which type of cluster should you identify?
A. Apache Storm B. Apache Hadoop C. Apache HBase D. Apache Spark
Answer: D
Question #13 (Topic: )
Note: This question is part of a series of questions that present the same Scenario.
Each question I the series contains a unique solution that might meet the stated
goals. Some question sets might have more than one correct solution while others
might not have correct solution.
You are implementing a batch processing solution by using Azure HDlnsight.
You have a data stored in Azure.
You need to ensure that you can access the data by using Azure Active Directory (Azure
AD) identities.
What should you do?
A. Use a shuffle join in an Apache Hive query that stores the data in a JSON format. B. Use a broadcast join in an Apache Hive query that stores the data in an ORC format. C. Increase the number of spark.executor.cores in an Apache Spark job that stores the data in a text format. D. Increase the number of spark.executor.instances in an Apache Spark job that stores the data in a text format. E. Decrease the level of parallelism in an Apache Spark job that Mores the data in a text format. F. Use an action in an Apache Oozie workflow that stores the data in a text format. Azure Data Factory linked service that stores the data in Azure Data lake. Azure DocumentDB database.
Answer: H
Question #14 (Topic: )
You have an Apache Spark cluster in Azure HDInsight. You execute the following
command,
%spark
import org.aache.spark.sql.hive.orc._
import org.apcahe.spark.sql._
What is the result of running the command?
A. the Hive ORC library is imported to Spark and external tables in ORC format are created. B. the Spark library is imported and the data is loaded to an Apache Hive table. C. the Hive ORC library is imported to Spark arid the ORC-formatted data stored in Apache Hive tables becomes accessible D. the Spark library is imported and Scala functions are executed
Answer: D
Question #15 (Topic: )
Note: This question is part of a series of questions that present the same Scenario.
Each question I the series contains a unique solution that might meet the stated
goals. Some question sets might have more than one correct solution while others
might not have correct solution.
Start of Repeated Scenario:
initial data that contains the crime data from major cities.
You plan to build training models from the training data. You plan to automate the process
of adding more data to the training models and to training the models by using the
additional data, including data that is collected in near real time. The system will be used to
analyze event data gathered from many different sources. Such as Internet of things (IoT)
increased crime risk at a particular time and ptace.
incoming data stream from
Facebook. which are event-based only, rather than time-based. You also have a time
interval stream every 10 seconds.
number that defines
how many times a hashtag occurs within a Facebook post or how many times a tweet that
contains a specific hashtag is retweeted.
You must use the appropriate data storage, stream analytics techniques, and Azure
HDInsight cluster types tor the various tasks associated to the processing pipeline.
End of repeated Scenario.
You are planning a storage strategy for a large amount of analytic data used for the crime
100 billion records, and more than
two billion records will be added daily.
You already created an Apache Hadoop cluster in HDInsight premium.
You need to implement the storage strategy to meet the following requirements:
The storage capacity must support 50 TB.
The storage must he optimized tor Hadoop.
The data must be stored in its native format
Enterprise-level security based on Active Directory must be supported.
What should you create?
Window, that has premium storage- a G-series size,
and uses Microsoft SQL Server 2016 to store the data
B. an Azure Data Lake Analytics service by using Azure Power Shell
C. an Azure Data Lake S
Answer: B
Download Exam
Page: 3 / 7
Total 35 questions