feat(sources): accept t.me links and auto-extract chat id #9

Merged
vmruiz merged 1 commit from feat/issue-4-tme-link into main 2026-08-23 19:03:23 +02:00
Owner

Summary

Accepts a t.me link in the Sources form and auto-extracts the numeric Telegram chat id, so sources can be created from a pasted link without leaving the value in a shape _resolve_chat cannot handle.

Closes #4

Problem

The "Telegram chat" field on the Sources form only accepted @username or a numeric chat id. A pasted https://t.me/somechannel was stored verbatim and a pasted https://t.me/+invitehash was stored verbatim. Both shapes later failed in _resolve_chat (app/telegram/indexer.py:47), which only understood @username or a digit-only id.

Approach

Add a single normaliser — _normalize_telegram_chat_identifier() in app/telegram/indexer.py — built on Telethon helpers:

Input Normalised form
@somechannel / somechannel @somechannel
https://t.me/somechannel / t.me/foo/ @somechannel
https://t.me/+HASH / t.me/+HASH numeric chat id (decoded via telethon.utils.resolve_invite_link)
https://t.me/joinchat/HASH numeric chat id
-100… / 123… unchanged (passthrough)
Anything else unchanged (let _resolve_chat raise with original text)
  • create_source (app/main.py) runs the field through the normaliser before storing it, so Source.telegram_chat is always a canonical shape.
  • _resolve_chat runs its input through the normaliser too, so sources stored with a t.me link in a prior release still resolve on the next indexing run.
  • The Sources form placeholder is updated to "@channel, chat id, or t.me link".

Tests

Added to tests/test_telegram_indexer.py:

  • Parameterised _normalize_telegram_chat_identifier cases for every accepted shape.
  • Whitespace-stripping test.
  • Negative tests for non-Telegram URLs and malformed invite hashes.
  • End-to-end _resolve_chat tests proving a t.me channel URL resolves via @username and a t.me/+HASH resolves via numeric id.

Sabotage check: replacing the normaliser with an identity function makes 12 of the new tests fail.

Verification

  • uv run pytest tests/test_telegram_indexer.py → 20 passed
  • uv run pytest --ignore=tests/test_job_recovery.py → 147 passed (3 excluded failures are pre-existing: they need a live Redis container and reproduce on clean origin/main)
  • uv run ruff check app/telegram/indexer.py app/main.py tests/test_telegram_indexer.py → no new issues

Risk / rollout

  • Stored value format change: bare usernames now stored as @somechannel. app/presentation.py:telegram_message_url already handles both shapes; existing _resolve_chat accepts both.
  • No migration needed; only newly-added sources go through the new path.
## Summary Accepts a `t.me` link in the Sources form and auto-extracts the numeric Telegram chat id, so sources can be created from a pasted link without leaving the value in a shape `_resolve_chat` cannot handle. Closes #4 ## Problem The "Telegram chat" field on the Sources form only accepted `@username` or a numeric chat id. A pasted `https://t.me/somechannel` was stored verbatim and a pasted `https://t.me/+invitehash` was stored verbatim. Both shapes later failed in `_resolve_chat` (`app/telegram/indexer.py:47`), which only understood `@username` or a digit-only id. ## Approach Add a single normaliser — `_normalize_telegram_chat_identifier()` in `app/telegram/indexer.py` — built on Telethon helpers: | Input | Normalised form | | --- | --- | | `@somechannel` / `somechannel` | `@somechannel` | | `https://t.me/somechannel` / `t.me/foo/` | `@somechannel` | | `https://t.me/+HASH` / `t.me/+HASH` | numeric chat id (decoded via `telethon.utils.resolve_invite_link`) | | `https://t.me/joinchat/HASH` | numeric chat id | | `-100…` / `123…` | unchanged (passthrough) | | Anything else | unchanged (let `_resolve_chat` raise with original text) | - `create_source` (`app/main.py`) runs the field through the normaliser before storing it, so `Source.telegram_chat` is always a canonical shape. - `_resolve_chat` runs its input through the normaliser too, so sources stored with a `t.me` link in a prior release still resolve on the next indexing run. - The Sources form placeholder is updated to "@channel, chat id, or t.me link". ## Tests Added to `tests/test_telegram_indexer.py`: - Parameterised `_normalize_telegram_chat_identifier` cases for every accepted shape. - Whitespace-stripping test. - Negative tests for non-Telegram URLs and malformed invite hashes. - End-to-end `_resolve_chat` tests proving a `t.me` channel URL resolves via `@username` and a `t.me/+HASH` resolves via numeric id. Sabotage check: replacing the normaliser with an identity function makes 12 of the new tests fail. ## Verification - `uv run pytest tests/test_telegram_indexer.py` → 20 passed - `uv run pytest --ignore=tests/test_job_recovery.py` → 147 passed (3 excluded failures are pre-existing: they need a live Redis container and reproduce on clean `origin/main`) - `uv run ruff check app/telegram/indexer.py app/main.py tests/test_telegram_indexer.py` → no new issues ## Risk / rollout - Stored value format change: bare usernames now stored as `@somechannel`. `app/presentation.py:telegram_message_url` already handles both shapes; existing `_resolve_chat` accepts both. - No migration needed; only newly-added sources go through the new path.
feat(sources): accept t.me links and auto-extract chat id
All checks were successful
Deploy Development Branch / deploy-dev (push) Successful in 49s
Clean Up Development Branch / cleanup-dev (pull_request) Successful in 35s
8364375f73
The Sources form previously only stored the Telegram chat field as
free text. Pasting a t.me/c/... channel link or a t.me/+... invite
link was stored verbatim and then failed in the indexer because
_resolve_chat only understood @username or a digit-only id.

Add _normalize_telegram_chat_identifier() in app/telegram/indexer.py:
  - https://t.me/somechannel / t.me/somechannel -> @somechannel
  - https://t.me/+HASH / t.me/joinchat/HASH -> numeric chat id
    (decoded via telethon.utils.resolve_invite_link)
  - numeric ids, @username, and bare usernames are preserved
    (bare usernames are normalised to the @-prefixed form)

create_source runs the field through the normaliser so the stored
Source.telegram_chat is always a canonical form that _resolve_chat
already understands. _resolve_chat also runs the input through the
normaliser so a previously-stored t.me link still resolves on the
next indexing run.

The sources.html placeholder now advertises 't.me link' as an
accepted shape.

Refs #4
vmruiz merged commit a07c9c1595 into main 2026-08-23 19:03:23 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
vmruiz/telegramarr!9
No description provided.