SimSms

1 Select service 1,296 services in stock

Choose a service

2 Select country 145 countries

Choose a country

3 Network & price

Pick a service and a country — in either order. The price and the network appear here.

Self-hosting an AI model: the accounts that still ask for a phone number

Open weights download in a single command and a GPU rents by the second — the compute stopped being the hard part. The friction is everything around it: the bot platform, the registry, the hub, each stopping to ask for a phone number at the worst possible moment. Which ones genuinely require one, and where a one-time code is the right answer.

A dotted path leaving a glowing GPU card, curving across the frame and ending inside a phone panel at a single orange dot

Running a language model on hardware you rent, from weights you downloaded yourself, is the part of this that has quietly become easy. One command pulls the model, a second starts the server, and the whole thing is billed by the second. What is not easy is everything around it: the hub that gates the download, the platform the model is supposed to answer on, the registry that holds your container, the dashboard that tells you the GPU is still alive. Several of those will stop and ask for a phone number, and they will ask at the least convenient moment.

This guide is about that gap. What a self-hosted setup actually costs to run, which of the surrounding accounts genuinely require a number and which only appear to, and where a one-time code is the right tool — and where it is the wrong one.

What self-hosting actually involves #

A self-hosted model is four moving parts, and only one of them is the model.

  • Compute — a GPU with enough memory to hold the weights, rented by the hour or the second. This is the only part billed continuously.
  • Weights — downloaded once from a model hub. Most open-weight models are a single sign-in away; a few sit behind an access request.
  • A surface — the thing people actually talk to: a Telegram bot, a Discord bot, a web endpoint, an internal API.
  • Plumbing — a container registry, somewhere to keep secrets, and something that watches the box.

The phone prompts never come from the model. They come from the surface and the plumbing, which are ordinary consumer and developer accounts with ordinary anti-abuse rules attached.

The compute is rented, and it is the cheap part #

Owning the card stopped making sense for most people a while ago. Hardware that can hold a 70B model unquantised costs more than a used car and sits idle most of the week; the same card rented by the second costs less than a takeaway for an evening of work. Per-second billing is the detail that matters — a job that takes four minutes should cost four minutes, not a rounded-up hour.

GPU What it comfortably holds On-demand, per hour
RTX 2070, 8 GB Speech-to-text, embeddings, small classifiers $0.032
RTX 4090, 24 GB A 7B–14B model at full speed, image generation $0.262
H100 SXM, 80 GB A 70B model unquantised, fine-tuning runs $1.512

Those are published on-demand rates at PowerGPU, which lists 79 NVIDIA models and sets each price against the public marketplace median instead of running an auction, so the figure on the page is the figure you pay. Billing is per second, payment is in crypto, and there is no identity check standing between you and the first instance. For a project you are deliberately keeping separate from your own name, that last point is the entire reason to look there rather than at a hyperscaler.

Which accounts ask for a number, and when #

The prompts are not evenly spread. Some platforms are phone-first and will not open an account at all without one; others only ask when something about the session looks unfamiliar. Working out which case you are in saves both money and a wasted number.

Account When the prompt appears Does a code settle it
Telegram, for a bot At account creation — there is no email-only path Yes. The clearest case on this list
Discord, for a bot On some sign-ups, and whenever an account gets flagged Yes — though the flag is usually what triggered it, not the sign-up
A model hub Rarely at sign-up; sometimes on gated models or organisation actions Usually yes, on the occasions it appears at all
A code host or registry At two-factor enrolment, where SMS is one option among several Possible, but an authenticator app is the better answer here
An AI API provider Once, at sign-up, before the first key is issued Yes, a single one-time code

Two patterns are worth separating, because they want different products. A number asked once, at sign-up, is a gate: you clear it and it does not come back. A number asked later, on a login the platform did not recognise, is a challenge: it can come back, and it will come back to the same number. Choosing as though a challenge were a gate is the most common and most expensive mistake in this whole process.

One-time code, or a number you keep #

If the account will only ever ask once, a one-time code is exactly right: take a number for the service you are on, the code arrives, the number goes back. If the account is something you will sign into again from new machines — a bot still running in six months — the number has to still exist when the challenge lands. That is what a rented number is for, and it is the difference between an account you keep and an account you quietly lose.

The order that works #

  1. 1 Decide the surface first Whether the model answers on Telegram, on Discord, or on a bare endpoint decides which accounts you need at all. A plain HTTP endpoint on a rented box needs no consumer account whatsoever — if that fits your project, you have just deleted every phone prompt on this page.
  2. 2 Create the accounts before you rent the GPU Accounts are the part that can stall for an hour. The GPU is billed from the second it starts. Do the stalling while nothing is running.
  3. 3 Order the number with the prompt already on screen Reach the field that is asking, leave it open, and only then order the code. A code is charged when it arrives, and it arrives whether or not you are ready for it. The Telegram numbers page and the Discord numbers page list the countries currently carrying each service.
  4. 4 Bind a second factor immediately The moment the account exists, add an authenticator app and save the recovery codes. That is what stops the platform ever needing to reach your number again — ninety seconds of work that removes a whole class of future problem.
  5. 5 Then start the instance Accounts settled, rent the GPU, pull the weights, start the server. From here the meter is running and nothing should be interrupting you.

What a virtual number does not do #

Worth being plain about the limits, because the failures are predictable ones.

  • It does not make an account anonymous on its own. The account still has an email, a payment trail and a behavioural fingerprint; the number is one identifier among several.
  • It does not get past a service that specifically rejects non-mobile ranges — see why services reject VoIP.
  • It does not undo a flag. If a platform has already decided an account looks wrong, a fresh code clears the challenge in front of you, not the reason it appeared.
  • It does not replace a second factor. A number is a fallback channel; an authenticator app is the thing that keeps the account yours.

If a code stops arriving mid-setup, the cause is almost never the model or the box — why your verification code never arrived covers the handful of reasons it actually happens.

The shape of a setup that holds #

The version of this that survives contact with reality looks much the same every time. Compute rented by the second and paid in crypto, so the bill matches the work and the instance can disappear when the project does. Weights held locally, so nobody can revoke access to your own model. A surface chosen to need as few consumer accounts as possible. And for the accounts you genuinely cannot avoid, a number that is there once for the gate, or still there in six months for the challenge — picked deliberately, not by whichever was cheapest on the day.

The compute has become the easy part. Plan the accounts with the same care you plan the GPU and the rest of it stops being interesting — which, for infrastructure, is the entire goal.

Read next