The author set up his own PII filter in front of cloud models and abandoned the self-hosted LLM. Over two months, he tested an RTX PRO 6000 96GB, RTX4090 48GB, and even an H200, evaluating options for self-hosting. He also bought an RTX5070ti and V100 32GB and ran many benchmarks. In the end, he opted for a personal data filter on his side rather than a model on his own hardware. The practical benefit is simple: to control sensitive data, you don’t have to host an LLM yourself — you can sanitize requests before sending them to the cloud.
Source: habr.com · post in Telegram