2026-08-21 21:08:19 : =================================================================================== 2026-08-21 21:08:19 : VSaaS Edge GPU recovery -- Fri Aug 21 09:08:19 PM UTC 2026 2026-08-21 21:08:19 : Bundle: /home/soterix/soterix_vsaas_edge_nvid 2026-08-21 21:08:19 : Logs dir: /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs 2026-08-21 21:08:19 : Run log: /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs/gpu_recovery_20260821_210819.log 2026-08-21 21:08:19 : Cmd trace: /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs/gpu_recovery_20260821_210819_trace.log 2026-08-21 21:08:19 : Diagnostics: /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs/diagnostics_20260821_210819 2026-08-21 21:08:19 : =================================================================================== 2026-08-21 21:08:19 : OS check passed: Ubuntu 22.04 2026-08-21 21:08:19 : #================================================== 2026-08-21 21:08:19 : # PHASE 1: Diagnostics 2026-08-21 21:08:19 : #-------------------------------------------------- 2026-08-21 21:08:19 : NVIDIA device(s) present on PCI bus: 2026-08-21 21:08:19 : 01:00.0 VGA compatible controller [0300]: NVIDIA Corporation Device [10de:27b2] (rev a1) 2026-08-21 21:08:19 : 01:00.1 Audio device [0403]: NVIDIA Corporation Device [10de:22bc] (rev a1) 2026-08-21 21:08:19 : Probing nvidia-smi (max 20s)... 2026-08-21 21:08:22 : nvidia-smi returned rc=6: No devices were found 2026-08-21 21:08:22 : nvidia-smi failed to talk to the driver. 2026-08-21 21:08:22 : NVRM Xid errors found in dmesg: 119  2026-08-21 21:08:22 :  Xid 119: GSP RPC TIMEOUT -- the GPU's GSP firmware has stopped answering 2026-08-21 21:08:22 :  the driver. Every nvidia-smi then blocks in the kernel forever. 2026-08-21 21:08:22 :  A warm reboot does NOT reset the GSP: the card keeps its state 2026-08-21 21:08:22 :  across it, so the fault survives and the box comes back broken. 2026-08-21 21:08:22 :  Only removing power resets the GSP. 2026-08-21 21:08:23 : Kernel module loaded in memory : 575.51.03 2026-08-21 21:08:23 : Kernel module on disk : 575.51.03 2026-08-21 21:08:23 : Target driver version : 575.51.03 2026-08-21 21:08:24 : dkms status: nvidia/575.51.03, 5.15.0-190-generic, x86_64: installed;nvidia/575.51.03, 6.8.0-136-generic, x86_64: installed;nvidia/575.51.03, 6.8.0-138-generic, x86_64: installed; 2026-08-21 21:08:24 : Running kernel: 6.8.0-138-generic 2026-08-21 21:08:24 : Device nodes present: /dev/nvidia0 /dev/nvidiactl /dev/nvidia-modeset /dev/nvidia-uvm /dev/nvidia-uvm-tools /dev/nvidia-caps: nvidia-cap1 nvidia-cap2 2026-08-21 21:08:24 : Loaded NVIDIA modules: 2026-08-21 21:08:24 : nvidia_uvm 2093056 0 2026-08-21 21:08:24 : nvidia_drm 139264 0 2026-08-21 21:08:24 : nvidia_modeset 1560576 1 nvidia_drm 2026-08-21 21:08:24 : nvidia 104640512 2 nvidia_uvm,nvidia_modeset 2026-08-21 21:08:24 : video 77824 1 nvidia_modeset 2026-08-21 21:08:24 : Processes stuck in uninterruptible (D) state -- unkillable, reboot likely required: 2026-08-21 21:08:24 :  887 D ext4lazyinit [ext4lazyinit] 2026-08-21 21:08:25 : Processes holding /dev/nvidia*: 0 line(s), see diagnostics. 2026-08-21 21:08:25 : Kernel taint flags: 12289 2026-08-21 21:08:25 : Hardware/PCIe/thermal errors present in dmesg -- see /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs/diagnostics_20260821_210819/dmesg_hw_errors.txt 2026-08-21 21:08:25 : /media/vdata is mounted. 2026-08-21 21:08:25 : /media/dbdata is mounted. 2026-08-21 21:08:25 : Docker daemon is running. 2026-08-21 21:08:25 : #================================================== 2026-08-21 21:08:25 : # Collecting detailed diagnostics 2026-08-21 21:08:25 : #-------------------------------------------------- 2026-08-21 21:08:25 : Target: /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs/diagnostics_20260821_210819 2026-08-21 21:08:25 : [gpu] 2026-08-21 21:08:26 : gpu/lspci_nvidia.txt 2026-08-21 21:08:28 : gpu/lspci_tree.txt 2026-08-21 21:08:29 : gpu/lspci_verbose.txt 2026-08-21 21:08:30 : gpu/dev_nodes.txt (rc=2) 2026-08-21 21:08:31 : gpu/proc_driver_version.txt 2026-08-21 21:08:32 : gpu/proc_gpus.txt 2026-08-21 21:08:33 : gpu/proc_params.txt 2026-08-21 21:08:34 : gpu/modprobe_conf.txt 2026-08-21 21:08:34 :  nvidia-smi collectors skipped (GPU wedged -- would leak D-state procs) 2026-08-21 21:08:34 : [kernel] 2026-08-21 21:08:35 : kernel/dmesg_full.txt 2026-08-21 21:08:36 : kernel/dmesg_nvidia.txt 2026-08-21 21:08:37 : kernel/xid_history.txt 2026-08-21 21:08:38 : kernel/dmesg_errors.txt 2026-08-21 21:08:39 : kernel/dmesg_hw.txt 2026-08-21 21:08:40 : kernel/uname.txt 2026-08-21 21:08:41 : kernel/cmdline.txt 2026-08-21 21:08:42 : kernel/tainted.txt 2026-08-21 21:08:43 : kernel/modules.txt 2026-08-21 21:08:44 : kernel/module_refcounts.txt 2026-08-21 21:08:45 : kernel/dkms_status.txt 2026-08-21 21:08:46 : kernel/headers.txt 2026-08-21 21:08:47 : kernel/secureboot.txt 2026-08-21 21:08:48 : kernel/interrupts.txt 2026-08-21 21:08:49 : kernel/iomem.txt 2026-08-21 21:08:49 : [system] 2026-08-21 21:08:50 : system/ps_full.txt 2026-08-21 21:08:51 : system/dstate_procs.txt 2026-08-21 21:08:52 : system/dstate_stacks.txt 2026-08-21 21:08:53 : system/gpu_holders.txt (rc=1) 2026-08-21 21:08:54 : system/uptime.txt 2026-08-21 21:08:55 : system/reboot_history.txt 2026-08-21 21:08:56 : system/meminfo.txt 2026-08-21 21:08:57 : system/df.txt 2026-08-21 21:08:58 : system/mounts.txt 2026-08-21 21:08:59 : system/sensors.txt 2026-08-21 21:09:00 : system/failed_units.txt 2026-08-21 21:09:01 : system/units_enabled.txt 2026-08-21 21:09:02 : system/crontab_root.txt 2026-08-21 21:09:02 : [journal] 2026-08-21 21:09:03 : journal/boots.txt 2026-08-21 21:09:04 : journal/this_boot_kernel.txt 2026-08-21 21:09:05 : journal/prev_boot_kernel.txt 2026-08-21 21:09:06 : journal/prev_boot_tail.txt 2026-08-21 21:09:07 : journal/this_boot_errors.txt 2026-08-21 21:09:09 : journal/docker.txt 2026-08-21 21:09:10 : journal/shutdown.txt 2026-08-21 21:09:10 : [docker] 2026-08-21 21:09:11 : docker/info.txt 2026-08-21 21:09:12 : docker/version.txt 2026-08-21 21:09:13 : docker/ps_all.txt 2026-08-21 21:09:14 : docker/images.txt 2026-08-21 21:09:15 : docker/daemon_json.txt 2026-08-21 21:09:16 : docker/runtimes.txt 2026-08-21 21:09:17 : docker/nvidia_ctk.txt 2026-08-21 21:09:18 : docker/compose_ps.txt 2026-08-21 21:09:19 : docker/compose_config.txt 2026-08-21 21:09:20 : docker/container_logs.txt 2026-08-21 21:09:20 : [packages] 2026-08-21 21:09:21 : packages/nvidia_cuda.txt 2026-08-21 21:09:22 : packages/holds.txt 2026-08-21 21:09:23 : packages/apt_history.txt 2026-08-21 21:09:24 : packages/unattended.txt 2026-08-21 21:09:25 : packages/sources.txt 2026-08-21 21:09:25 : [bundle] 2026-08-21 21:09:26 : bundle/listing.txt 2026-08-21 21:09:27 : bundle/engines.txt 2026-08-21 21:09:28 : bundle/pro_cnf.txt 2026-08-21 21:09:29 : bundle/recovery_state.txt (rc=1) 2026-08-21 21:09:29 :  Detailed diagnostics collected under /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs/diagnostics_20260821_210819 2026-08-21 21:09:29 : #================================================== 2026-08-21 21:09:29 : # Assessment 2026-08-21 21:09:29 : #-------------------------------------------------- 2026-08-21 21:09:29 : LIKELY CAUSE: The GPU's GSP firmware stopped responding (Xid 119/120). Every 2026-08-21 21:09:29 :  driver call then blocks in the kernel forever, and the tasks that 2026-08-21 21:09:29 :  made them become unkillable -- which also stops the box shutting 2026-08-21 21:09:29 :  down cleanly. 2026-08-21 21:09:29 : ACTION: Cold power cycle (mains off ~30s). A warm reboot does NOT reset the GSP, 2026-08-21 21:09:29 :  so the fault survives it. To stop this recurring, re-run with 2026-08-21 21:09:29 :  --disable-gsp, then power cycle. 2026-08-21 21:09:29 : #================================================== 2026-08-21 21:09:29 : # PHASE 2: Repair 2026-08-21 21:09:29 : #-------------------------------------------------- 2026-08-21 21:09:29 : The GPU's GSP firmware has stopped responding (Xid 119/120). 2026-08-21 21:09:29 : No software repair can reach the card while its firmware is wedged. 2026-08-21 21:09:29 : This is NOT recoverable by a warm reboot -- a full power cycle is required. 2026-08-21 21:09:29 : #================================================== 2026-08-21 21:09:29 : # Recovery summary 2026-08-21 21:09:29 : #-------------------------------------------------- 2026-08-21 21:09:32 : nvidia-smi returned rc=6: No devices were found 2026-08-21 21:09:32 : GPU state : error 2026-08-21 21:09:32 : Diagnostics saved : /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs/diagnostics_20260821_210819 2026-08-21 21:09:32 : Full log : /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs/gpu_recovery_20260821_210819.log 2026-08-21 21:09:32 : GPU FIRMWARE (GSP) FAULT -- Xid 119/120. 2026-08-21 21:09:32 : Power the machine off completely (mains off ~30 seconds), then power on. 2026-08-21 21:09:32 : A warm reboot will not clear this: the GSP keeps its state across one, 2026-08-21 21:09:32 : so the box comes back with the same fault. That matches what you have 2026-08-21 21:09:32 : been seeing -- only the power cycle works. 2026-08-21 21:09:32 :  2026-08-21 21:09:32 : To stop it recurring, run the GPU with GSP firmware disabled: 2026-08-21 21:09:32 :  sudo bash recover_nexaiq_edge_gpu.sh --disable-gsp # then power cycle 2026-08-21 21:09:32 : This sets NVreg_EnableGpuFirmware=0, which is the standard workaround 2026-08-21 21:09:32 : for repeated Xid 119 on datacentre/RTX cards. 2026-08-21 21:09:32 : Send /home/soterix/soterix_vsaas_edge_nvid/logs/recoverylogs/diagnostics_20260821_210819 to support.