Unsolved
4 Posts
0
162
March 20th, 2024 15:08
Problems with VMs and VMware after error on SSD that is in Raid-1
Good morning everyone, I hope you are all doing well! I have the Dell PowerEdge R750xs server, which after setting up a Raid 10 using the PERC H355 Front (Embedded) controller, started experiencing issues with VMware and its VMs. However, this VMware is on another Datastore with Raid-1, and in the logs, it was found that Disk-3 of the Raid 1 where the VMs are located is having problems. At the moment, this raid-1 is operating with only one functional disk.
Additionally, a critical fault was detected on drive 3 in disk drive bay 1 at Wed Mar 20 2024 07:33:32.


No Events found!


Origin3k
6 Operator
•
2.4K Posts
•
12.4K Points
0
March 20th, 2024 16:47
The Disk3 is dead or the slot is faulty. The later one is unlikely. If a disk died which was a member of a Disk group the VD(s) are going into a degraded mode.
I find it annoying that MediaType shows "HDD" because all other are SSDs and within a Raidgroup all disks have to be the same MediaType.
Regards,
Joerg
Carlos Eduardo Pereira
4 Posts
0
March 20th, 2024 17:02
@ Origem3k Alright, at the moment I'm proceeding with the RAID 1 containing only one disk, with the faulty one removed. If I migrate all VMs to the RAID 10 and leave only the VMware system on this RAID-1 until maintenance, will it continue to function properly? Or am I at risk of encountering issues with this server due to this setup?
Origin3k
6 Operator
•
2.4K Posts
•
12.4K Points
0
March 20th, 2024 17:17
From the Screenshots i cant see which VD is on which Raidgroup. I assume that Disk 0.1.2 and 0.1.3 are one of your two RAID1 Diskgroups? So if you have VMs on a ~920GB ESXi Datastore you should migrate these with a svMotion to another Datastore and after Replacing the failed SSD in Disk3 you need to verifying of the Rebuild was successfully and than you can migrate VMs back.
If the replaced Disk not automaticly went into the diskgroup and rebuild is startet you should mark these as a Hotspare and assign the Hostspare to the second RAID1 Group.
Regards,
Joerg
Carlos Eduardo Pereira
4 Posts
0
March 20th, 2024 17:46
@ Origem3k Yes, this RAID-1 consists of disks 0:1:2 and 0:1:3. Disk 0:1:3 indeed failed. One last question, if I keep disk 0:1:3 in the server until tonight but have only disk 0:1:2 functioning in the RAID, am I at risk with active services? They are being maintained on another 1TB SSD. Or is it possible to disable disk 0:1:3 through iDRAC until I remove it?
Thank you very much for the help.
Origin3k
6 Operator
•
2.4K Posts
•
12.4K Points
0
March 20th, 2024 18:14
If you loose also Disk 0:1:2 your RAID1 is gone and the ESXi Datastore disapears and your VMs went away.
Next time think about a (Global)Hotspare or a different RAID Level which support higher redundancy.
Is ESXi installed on the 2x480GB (RAID1)? Next time buy 4x 960 and create a RAID5+HS or better a RAID6 and than create 2 VDs on that Diskgroup. A small one for ESXI(128GB) and the second as the first Datastore. As long as you never want to increase such a setup it will get better redundancy and capacity for a little bit more $$$.
Regards,
Joerg
Carlos Eduardo Pereira
4 Posts
0
March 20th, 2024 19:22
@ Origem3k
Alright, to replace it, I set up a RAID 10 with 4x4TB drives. Thank you very much!
Origin3k
6 Operator
•
2.4K Posts
•
12.4K Points
0
March 20th, 2024 19:43
A RAID10 is also a single Redundany one.. only with a lot of luck you can survive 2 disk failures with a R10.
Regards,
Joerg