Networking

Sweet Spot on a Cheap Cloud GPU: MoE, Vision Ceilings, and VRAM Budgeting

Sweet Spot on a Cheap Cloud GPU: MoE, Vision Ceilings, and VRAM Budgeting

117 tokens per second from a “35-billion-parameter” model on a $0.27/hr GPU. The counterintuitive part: a smaller 27B model on the same GPU manages 33 tok/s — about 3.5× slower. The difference is not the parameter count on the spec sheet. It is how many of those parameters actually fire on every token.

This is the story of an afternoon spent learning to read a VRAM budget — and why “it fits” is the wrong question to ask.

From Local to Cloud GPU: Ollama on a Rented A100

From Local to Cloud GPU: Ollama on a Rented A100

134 tokens per second on a 35-billion-parameter model. That is what an NVIDIA A100 produces for $1.40 an hour — rented in the cloud, no hardware to buy, no drivers to fight. The plan was simple: spin up a pod, install Ollama, and see what consumer-unfriendly hardware actually buys you.

The GPU vanished the next morning. And the cheaper replacement taught me something the A100 couldn’t.

Lubuntu ssh: Read from socket failed – solved

SSH to my fresh lubuntu on cubieboard2 was not working. I tried it also from the lubuntu itself with the error:

$ ssh localhost
Read from socket failed: Connection reset by peer

I went through many google search results. But I could not find any appropriate solution. Finally did what I should as the first step. Uninstall and install the ssh.

This page helped me to try uninstall with complete removal (also the configuration files). It should work with the command