Hi friends. Yes, I know. 4gb vram is too small for use or try to inference in some decents models. Buy it’s what i have right now. Until i can bought some better gpu. But, anyway, i want to experimet with that. Wich models for tts, some coding assist, or for scripting in my local machine with linux . Can you help me with some little models names? Thank you very much!
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Recommend Compute for running a TensorRT-LLM using LLama2 13B & 70B model | 2 | 1219 | November 15, 2023 | |
| How to add custom model to chat with rtx? | 6 | 7704 | February 23, 2024 | |
| Turbocharging Meta Llama 3 Performance with NVIDIA TensorRT-LLM and NVIDIA Triton Inference Server | 61 | 5066 | August 28, 2024 | |
| Optimizing Inference on Large Language Models with NVIDIA TensorRT-LLM, Now Publicly Available | 8 | 2194 | January 25, 2024 | |
| NVIDIA TensorRT-LLM 및 NVIDIA Triton Inference Server로 Meta Llama 3 성능 강화 | 0 | 415 | May 3, 2024 | |
| Supercharging Llama 3.1 across NVIDIA Platforms | 13 | 598 | September 17, 2024 | |
| Deploying GPT-J and T5 with FasterTransformer and Triton Inference Server | 7 | 1224 | April 19, 2023 | |
| NVIDIA folks -- where is this promised nvfp4 speedup? | 27 | 3346 | March 26, 2026 | |
| Better GPU for training & Inference & Execution LLModels | 1 | 652 | November 30, 2023 | |
| ChatRTX compatibility with A100 GPU | 0 | 357 | August 6, 2024 |