Every hour, a fresh AI news story goes out to the @ai_release1 channel; once a day — a digest; every 8 hours — a video; plus a daily Windows tip and a weekly "roundup" of related topics. All of it runs on a single budget VPS with 1 GB of RAM, with no human involvement and no desktop PC. Over a month we've accumulated ~450 posts and 716 subscribers. This article covers how the pipeline is built, what pitfalls we hit, and how we save tokens.
The original idea was simple: stop reading AI news manually and instead get it in Telegram in a condensed, verified form. But it quickly became clear that "just RSS + a summary" doesn't work:
So the stack ended up "thin but mean": Python scripts on systemd timers, external LLM APIs, lightweight metasearch, and a static mirror site.
The project is deliberately kept to a minimal sum:
| Expense | Cost |
|---|---|
| VPS (1 GB RAM) | ~$5/mo |
| OpenCode Zen GO subscription (LLM API) | $10/mo |
| Domain ai-release.net | $22/yr (~$1.8/mo) |
| Telegram Premium | $25/yr (~$2/mo) |
| **Total** | **~$18.8/mo** |
The most expensive item is the LLM subscription, but it covers news, the chatbot, and editing. Cover images cost nothing — we assemble them from open sources (see below).
Everything rests on systemd timers and a handful of Python scripts:
ai-release-hourly.timer -> hourly.py (*:15) 1 news item per hour
ai-release-video.timer -> video.py (0:30, 8:30, 16:30) 1 video every 8 hours
ai-release-daily.timer -> daily.py (15:00 UTC / 18:00 MSK) daily digest
ai-release-stats.timer -> stats.py (18:00 UTC) statistics
ai-release-chat.service -> chat.py (daemon) AI replies in DMs
ai-release-web.service -> site.py / web.py / guides.py web mirror
The schedule is set via `OnCalendar`, for example for the hourly timer:
[Timer]
OnCalendar=*-*-* *:15:00
Persistent=true
No message broker, no "orchestrator" — just timers. For a project of this size, that's the right call: each run is a separate process that can crash without taking down the rest.
Publishing goes directly through the Telegram Bot API. News sources: RSS aggregators (Habr, Ars Technica, TechCrunch, VentureBeat, The Register, The Decoder, and others), primary sources (the OpenAI blog, Google Research, the ArXiv API, Hugging Face Daily Papers, GitHub Trending, Hacker News via the Algolia API), and author blogs (Simon Willison, Karpathy, Sebastian Raschka, Nathan Lambert). Backup search — a local SearXNG in Docker.
The first version published everything and turned into a feed of duplicates. Now there are three levels of filtering:
1. **Scoring.** Each news item gets a score based on freshness, source weight, and headline quality. Primary sources and authors weigh more than aggregators; fresher beats older.
2. **Headline hashing.** Before comparison, the headline is normalized: publisher suffixes are trimmed (`— TechCrunch`, `| WindowsLatest`), leading `BREAKING:`/`EXCLUSIVE` tags are stripped, trailing phrases like `report says` are removed, and quotes and dashes are unified. Then it's hashed and checked against SQLite:
def normalize_title(t: str) -> str:
t = re.sub(r"\s*[-—|]\s*(TechCrunch|Engadget|The Verge|WindowsLatest).*$", "", t, flags=re.I)
t = re.sub(r"^(BREAKING|EXCLUSIVE)\s*:?\s*", "", t, flags=re.I)
t = t.replace("«", '"').replace("»", '"').strip()
return t.lower()
3. **Semantic dedup via LLM.** Even with normalization, two sources can describe the same event in different words. So the candidate is additionally compared against the text of recent channel posts, and if it's "the same event" — it's dropped.
There's also a separate SEO-junk filter: domains not on the allowlist only pass through with a specific signal (a model, company, or event name), not with broad terms like "ai/neural network." Plus a blacklist of promotional markers ("magazine about," "fresh issue," "subscribe").
This turned out to be the most important part. Previously, summaries would sometimes present someone else's conclusion as fact and invent causes. We tightened the prompt: state **only** what's in the headline and description; don't invent causes, numbers, or consequences; preserve attribution ("according to the study," "journalists believe") when a conclusion belongs to the source. We also lowered `temperature` from 0.8 to 0.5.
A real example: a post about Windows 10 presented the publication's opinion as fact and invented causes. We had to introduce a separate factual rule — "Windows 10 support ended on 2025-10-14, there are no free security updates" — so the editor adds security context whenever a story advises staying on the old OS.
The LLM is the most expensive part of the pipeline, so we:
For long summaries, the batch is built as a single request with headroom in `max_tokens` — otherwise the model spends tokens on reasoning and cuts the answer short.
Image generation ran into a paid model and limits. So the cover is assembled "for free," in several steps:
1. the source article's `og:image`;
2. a stock image on the topic via metasearch (allowlist: Unsplash, Pexels, Pixabay, Wikimedia, Flickr, pinimg);
3. a themed anime/manga illustration;
4. a pre-made fallback set.
Images are additionally checked with a HEAD request. If Telegram can't fetch the image by URL (HTTP 400), the post isn't lost — another cover is substituted and the publish is retried.
Networks and external APIs go down regularly. What we did:
Each post has a page on the static mirror site `ai-release.net` (RU + EN), plus a sitemap with `hreflang`, `robots.txt`, and instant indexing via **IndexNow** (Bing + Yandex) on every publish. Analytics — Google Analytics 4 and Yandex.Metrika.
A separate topic is **AI SEO**, so that neural networks and AI aggregators pick up the content themselves:
`robots.txt` is open to all AI bots (GPTBot, PerplexityBot, ClaudeBot, Google-Extended). Over a month, this produced a noticeable share of traffic from LLM crawlers.
The bot replies to subscribers in DMs through the same LLM. A public chat is untrusted input, so protection is built in three layers:
Dialogue history is also capped in length so the context doesn't grow and drag leaks along with it.
1. A "smart" feed doesn't need a server with a GPU — a well-designed pipeline and external LLMs are enough.
2. Deduplication and fighting hallucinations matter more than model quality.
3. Token savings are primarily about caching and routing, not "a cheaper model."
4. For indexing in the LLM era, you need a separate layer: `llms.txt`, markdown versions, and open crawlers.
---
AI news every hour and deep dives — in the @ai_release1 channel; the post archive and guides — on ai-release.net.
💬 **Questions and discussion — in the comments below.**