# Deploying GPT-J and T5 with FasterTransformer and Triton Inference Server

**URL:** <https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838>\
**Category:** Technical Blog\
**Created:** [August 3, 2022, 5:00pm UTC](https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838 "2022-08-03T17:00:23Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![jwitsoe](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/jwitsoe/32/16591_2.png) [@jwitsoe](https://forums.developer.nvidia.com/u/jwitsoe)\
**Post date:** [August 3, 2022, 5:00pm UTC](https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838/1 "2022-08-03T17:00:23Z")

</div>

Originally published at: [Deploying GPT-J and T5 with NVIDIA Triton Inference Server | NVIDIA Technical Blog](https://developer.nvidia.com/blog/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/)  
   
Learn step by step how to use the FasterTransformer library and Triton Inference Server to serve T5-3B and GPT-J 6B models in an optimal manner with tensor parallelism.

---

<div class="post-metadata">

**Author:** ![kun.he.love.u](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@kun.he.love.u](https://forums.developer.nvidia.com/u/kun.he.love.u)\
**Post date:** [August 12, 2022, 12:21am UTC](https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838/2 "2022-08-12T00:21:34Z")

</div>

hi @jwitsoe ,  
I am from the Chinese developer community.

There seems to be a picture mismatch in the results section of the article. Figure 5 should be T5-3B model inference speed-up comparison, but it shows GPT-J 6B.

 ![微信截图_20220812082056](https://global.discourse-cdn.com/nvidia/original/3X/f/1/f134ea747bf126495fcb1b37a94213ce7df8da50.png)

---

<div class="post-metadata">

**Author:** ![dtimonin](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@dtimonin](https://forums.developer.nvidia.com/u/dtimonin)\
**Post date:** [August 15, 2022, 5:26pm UTC](https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838/3 "2022-08-15T17:26:44Z")

</div>

@kun.he.love.u Thanks for letting us know. We’ve updated the image.

---

<div class="post-metadata">

**Author:** ![aneesinaec](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@aneesinaec](https://forums.developer.nvidia.com/u/aneesinaec)\
**Post date:** [April 3, 2023, 10:09am UTC](https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838/4 "2023-04-03T10:09:33Z")

</div>

Can you please provide similar step by step guide for multi node inference example with Triton server?

---

<div class="post-metadata">

**Author:** ![jarretbembry](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/jarretbembry/32/241360_2.png) [@jarretbembry](https://forums.developer.nvidia.com/u/jarretbembry)\
**Post date:** [April 6, 2023, 5:58am UTC](https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838/5 "2023-04-06T05:58:23Z")

</div>

wow, I was waiting for such a guide for a long. Waiting for smth like that with Trition server as was asked above by aneesinaec.

---

<div class="post-metadata">

**Author:** ![aneesinaec](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@aneesinaec](https://forums.developer.nvidia.com/u/aneesinaec)\
**Post date:** [April 17, 2023, 10:01am UTC](https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838/6 "2023-04-17T10:01:48Z")

</div>

I tried similar exercise with a bloom model.  
I have 2 GPU’s with ~10 GB Memory on each. While trying to load a 14GB model in 2 GPU config, I keep getting out of memory error.

Fasttransformer backend supposed to split the 14 GB model in to 2 CPU’s and load…rite? What am i possibly missing here?

---

<div class="post-metadata">

**Author:** ![bhsueh](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@bhsueh](https://forums.developer.nvidia.com/u/bhsueh)\
**Post date:** [April 19, 2023, 1:11am UTC](https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838/7 "2023-04-19T01:11:25Z")

</div>

Thank you for the feedback. We will consider it. You can also refer the document in [fastertransformer\_backend/t5\_guide.md at main · triton-inference-server/fastertransformer\_backend · GitHub](https://github.com/triton-inference-server/fastertransformer_backend/blob/main/docs/t5_guide.md) first.

---

<div class="post-metadata">

**Author:** ![bhsueh](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@bhsueh](https://forums.developer.nvidia.com/u/bhsueh)\
**Post date:** [April 19, 2023, 1:12am UTC](https://forums.developer.nvidia.com/t/deploying-gpt-j-and-t5-with-fastertransformer-and-triton-inference-server/222838/8 "2023-04-19T01:12:25Z")

</div>

Can you share more details and your steps?
