By vilius
Hello,
AIX 6.1 TL7 SP6
POwerHA 6.1 SP10
I was experimenting with new hacmp build. It’s 3-node cluster build on AIX 6.1 lpars. It contains Ethernet and diskhb networks. Shared vg disk is SAN disk. Two nodes see disk using vscsi, third node sees disk using npiv. Application is db2 server.
Most accidents usually involve some kind of network failure – so I decided to test my cluster against Ethernet failure and SAN failure. Ethernet failure test was successful – when node lost Ethernet connectivity(both cables of course) my resource group jumped to next node with no problem.
Next I did SAN failure test:
I did it in 2 different ways by removing vsci mapping in vios or by removing fcs mapping in vios(npiv case) – results were exactly the same in both cases – cluster reacted correctly and started release_vg_fs event, release_vg_fs script tried to unmount filesystems but since all fs disk devices were gone script just hung, and cluster started issuing config_too_long events..
So clstat reports resouce group as “RELEASING..” and that’s it…
How do I configure PowerHA to handle full vg loss(for example SAN down causes that) correctly ??
thanks,
Vilius M.
…read more
Source: FULL ARTICLE at The UNIX and Linux Forums