2016-01-13 96 views
3

我在這些instructions之後的virtualenv中安裝了tensorflow的GPU版本。問題是,開始會話時出現分段錯誤。也就是說,該代碼:在virtualenv上運行GPU集羣上的tensorflow

import tensorflow as tf 
sess = tf.InteractiveSession() 

退出並出現以下錯誤:

(tesnsorflowenv)[email protected]$ python testtensorflow.py 
I tensorflow/stream_executor/dso_loader.cc:101] successfully opened CUDA library libcublas.so.7.0 locally 
I tensorflow/stream_executor/dso_loader.cc:93] Couldn't open CUDA library libcudnn.so.6.5. LD_LIBRARY_PATH: :/vol/cuda/7.0.28/lib64 
I tensorflow/stream_executor/cuda/cuda_dnn.cc:1382] Unable to load cuDNN DSO 
I tensorflow/stream_executor/dso_loader.cc:101] successfully opened CUDA library libcufft.so.7.0 locally 
I tensorflow/stream_executor/dso_loader.cc:101] successfully opened CUDA library libcuda.so locally 
I tensorflow/stream_executor/dso_loader.cc:101] successfully opened CUDA library libcurand.so.7.0 locally 
I tensorflow/core/common_runtime/local_device.cc:40] Local device intra op parallelism threads: 40 
Segmentation fault 

我嘗試使用gdb的深入挖掘,但只得到了以下額外產出:

[New Thread 0x7fffdf880700 (LWP 32641)] 
[New Thread 0x7fffdf07f700 (LWP 32642)] 
... lines omitted 
[New Thread 0x7fffadffb700 (LWP 32681)] 
[Thread 0x7fffadffb700 (LWP 32681) exited] 
Program received signal SIGSEGV, Segmentation fault. 
0x0000000000000000 in ??() 

任何想法這裏發生了什麼以及如何解決它?

這裏是NVIDIA-SMI的輸出:

+------------------------------------------------------+      
| NVIDIA-SMI 352.63  Driver Version: 352.63   |      
|-------------------------------+----------------------+----------------------+ 
| GPU Name  Persistence-M| Bus-Id  Disp.A | Volatile Uncorr. ECC | 
| Fan Temp Perf Pwr:Usage/Cap|   Memory-Usage | GPU-Util Compute M. | 
|===============================+======================+======================| 
| 0 Tesla K80   On | 0000:06:00.0  Off |     0 | 
| N/A 65C P0 142W/149W | 235MiB/11519MiB |  81% E. Process | 
+-------------------------------+----------------------+----------------------+ 
| 1 Tesla K80   On | 0000:07:00.0  Off |     0 | 
| N/A 25C P8 30W/149W |  55MiB/11519MiB |  0% E. Process | 
+-------------------------------+----------------------+----------------------+ 
| 2 Tesla K80   On | 0000:0D:00.0  Off |     0 | 
| N/A 27C P8 26W/149W |  55MiB/11519MiB |  0% Prohibited | 
+-------------------------------+----------------------+----------------------+ 
| 3 Tesla K80   On | 0000:0E:00.0  Off |     0 | 
| N/A 25C P8 28W/149W |  55MiB/11519MiB |  0% E. Process | 
+-------------------------------+----------------------+----------------------+ 
| 4 Tesla K80   On | 0000:86:00.0  Off |     0 | 
| N/A 46C P0 85W/149W | 206MiB/11519MiB |  97% E. Process | 
+-------------------------------+----------------------+----------------------+ 
| 5 Tesla K80   On | 0000:87:00.0  Off |     0 | 
| N/A 27C P8 29W/149W |  55MiB/11519MiB |  0% E. Process | 
+-------------------------------+----------------------+----------------------+ 
| 6 Tesla K80   On | 0000:8D:00.0  Off |     0 | 
| N/A 28C P8 26W/149W |  55MiB/11519MiB |  0% Prohibited | 
+-------------------------------+----------------------+----------------------+ 
| 7 Tesla K80   On | 0000:8E:00.0  Off |     0 | 
| N/A 23C P8 30W/149W |  55MiB/11519MiB |  0% E. Process | 
+-------------------------------+----------------------+----------------------+ 

感謝在這個問題上的任何幫助!

+0

請嘗試從源代碼構建以下說明[這裏](https://www.tensorflow.org/versions /master/get_started/os_setup.html#installing-from-sources),最好以調試模式運行,並提供完整的堆棧跟蹤。這可能有助於查明SIGSEGV的來源。 – keveman

回答

4

它沒有找到CuDNN -

I tensorflow/stream_executor/dso_loader.cc:93] Couldn't open CUDA library > libcudnn.so.6.5. LD_LIBRARY_PATH: :/vol/cuda/7.0.28/lib64 I tensorflow/stream_executor/cuda/cuda_dnn.cc:1382] Unable to load cuDNN DSO

你需要把它安裝。請參閱the TensorFlow CUDA installation instructions

+0

是的,可能就是這樣!不知何故,當我在本地機器上測試時,這個問題沒有顯現出來。謝謝你的幫助。在我的cudnn應用程序獲得批准後,我會知道它是否工作... – Chrigi

0

後解壓cudnn

[[email protected] cudnn]# cd include/ 
[[email protected] include]# mv cudnn.h /usr/local/cuda/include/ 
[[email protected] include]# cd ../lib64/ 
[[email protected] lib64]# mv * /usr/local/cuda/lib 

而且它是確定

[[email protected] ~]# python 
Python 2.7.5 (default, Sep 15 2016, 22:37:39) 
[GCC 4.8.5 20150623 (Red Hat 4.8.5-4)] on linux2 
Type "help", "copyright", "credits" or "license" for more information. 
>>> import tensorflow as f 
I tensorflow/stream_executor/dso_loader.cc:128] successfully opened CUDA library libcublas.so.8.0 locally 
I tensorflow/stream_executor/dso_loader.cc:128] successfully opened CUDA library libcudnn.so.5 locally 
I tensorflow/stream_executor/dso_loader.cc:128] successfully opened CUDA library libcufft.so.8.0 locally 
I tensorflow/stream_executor/dso_loader.cc:128] successfully opened CUDA library libcuda.so.1 locally 
I tensorflow/stream_executor/dso_loader.cc:128] successfully opened CUDA library libcurand.so.8.0 locally 
>>>