NVIDIA NCP-AII - NCP-AI Infrastructure Exam

Question #11 (Topic: Exam A)
A systems engineer is updating firmware across a large DGX cluster using automation.
What is the best practice for minimizing risk and ensuring cluster health during and after the process?
A. To save time, simultaneously update all nodes in the cluster without draining or diagnostics. B. Drain nodes from the scheduler, run pre-update diagnostics, update firmware in batches, and verify health post-update before scaling to the next batch. C. Drain nodes from the scheduler, update firmware in batches, skip diagnostics and verify health post-update before scaling to the next batch. D. Update nodes that have reported faults, leaving others on older firmware.
Answer: B
Question #12 (Topic: Exam A)
A customer has just completed the first boot of their DGX system and is prompted to create an administrative user.
What is the correct approach for setting up this user to ensure secure BMC and GRUB access?
A. Create a unique, strong, lower-case username and password that will be used for both BMC and GRUB access, avoiding default or weak credentials. B. Skip the creation of a new user and retain the default admin account for BMC and GRUB access. C. Create separate usernames for BMC and GRUB to maximize flexibility. D. Use “sysadmin” as the username and a simple password for ease of management.
Answer: A
Question #13 (Topic: Exam A)
As the infrastructure lead for an NVIDIA AI Factory deployment, you have just uploaded the latest supported firmware packages to your DGX system. It is now critical to ensure all hardware components run the new firmware and the DGX returns to full operational capability. Which sequence best guarantees that all relevant components are correctly running updated firmware according to NVIDIA’s documentation and recommended operational steps?
A. Execute a single AC power cycle on the DGX after the update process, then reset the software stack and verify status using diagnostic commands on each node for confirmation of all component updates. B. Initiate a cold power cycle on the system to activate firmware for components, reset the BMC using the recommended command, and perform an AC power cycle to ensure EROT and CPLD firmware is activated. C. Perform a software-driven restart on the operating system of every compute node, then use advanced tools to check firmware status, and reissue update commands if any firmware appears inactive afterward. D. Initiate a cold power cycle on all node trays to activate firmware, follow with a DGX reboot procedure, and use the management interface to finish activating CPLD firmware on the host.
Answer: B
Question #14 (Topic: Exam A)
You are tasked with updating both NVIDIA GPU drivers and DOCA drivers on a set of servers used for AI workloads. The environment previously had an older driver stack and custom kernel modules.
What is the most important step to successfully upgrade the drivers without causing conflicts?
A. Update the GPU driver leaving the DOCA and OFED drivers unchanged as long as they are detecting the hardware properly. B. Uninstall all existing GPU and DOCA-related drivers and associated kernel modules before the new install C. Keep the older driver running alongside the new version in case you need to roll back the upgrade. D. Validate the driver version post-install since the fresh install will overwrite the legacy drivers.
Answer: B
Question #15 (Topic: Exam A)
During multi-node HPL burn-in, GPUs show uneven utilization.
Which configuration ensures balanced workload distribution?
A. HPL_RUN_GEMM_TESTS to skip validation B. Set --gpu-affinity and --cpu-affinity to align GPU and NUMA nodes C. HPL_OOC_TILE_M to 8192 for larger blocks D. Enable HPL_USE_NVSHMEM=1 for shared memory acceleration
Answer: B
Download Exam
Page: 3 / 34
Total 170 questions