My 2 pixels:
For 24/7 over a long timespan, the F@H folks have quite some experience. Well, their cluster is based on R580s from AMD, but they’re quite content in terms of MTBF. Some language about this has been published here:
@article{Owens:2008:GC,
title = "GPU Computing",
journal = "Proceedings of the IEEE",
author = "John D. Owens AND Mike Houston AND David Luebke AND Simon Green AND John E. Stone AND James C. Phillips ",
year = "2008",
month = may,
keywords = "GPGPU, GPU computing, parallel computing",
pages = "879--899",
volume = "96",
number = "5"
}
My own cluster (4 8800 GTX’s and 8 dualcores in four nodes) is not used 24/7, but it has been running without any HW failures for more than a year now.
Another cluster I have worked on used >100 Quadro 1400s for more than 2.5 years now, and in that time, one GPU died. For reference: Almost 25% of the DDR RAM died in the same time…
For the architecture-aware readers, this paper is quite interesting:
@INPROCEEDINGS{Sheaffer:2007:AHR,
author = {Jeremy W. Sheaffer and David P. Luebke and Kevin Skadron},
title = {A Hardware Redundancy and Recovery Mechanism for Reliable Scientific Computation on Graphics Processors},
booktitle = {Graphics Hardware 2007},
year = {2007},
editor = {Timo Aila and Mark Segal},
month = aug,
pages = {55--64},
}