Data collectors - how to collect data continuously
A collector retrieves data from a source without manual work. How to build one so it does not fail silently, and why failure handling takes most of the work.

A data collector is a process that periodically retrieves data from a source, validates it, and saves it to your database — without manually starting each run. It is an important part of the data layer for dashboards and forecasts that need regular updates; one-off analysis may use a controlled export.
An estimate must include more than retrieval: validation, duplicates, retries, delayed data, monitoring, retention, and recovery after failure. For a simple source, this surrounding work can take more effort than the API call itself.
What a solid collector consists of
Match these six elements to the source, required freshness, and cost of a gap:
- Retrieval from the source — API, database, file, or website.
- Validation — whether data has expected structure and sensible values. A collector that saves everything it receives silently poisons the database.
- Timestamp and provenance — when it was retrieved and from where. Without them, discrepancies cannot be diagnosed.
- Idempotent writing — repeating the same retrieval must not duplicate data. This makes failed runs safe to retry.
- Error handling with retries — with increasing delay, so a troubled source is not overwhelmed.
- Alert and log — a signal that something is wrong before someone notices it in a report.
An alert is not a substitute for completeness checks. You also need an owner, expected data-delivery time, and a gap-filling procedure. Incomplete data without a clear marker can lead to wrong decisions.
Four retrieval methods and when to use each
| Method | How it works | When to use it |
|---|---|---|
| Periodic polling | ask the source for new data every X minutes | when an API supports incremental reads and acceptable delay fits the interval |
| Incoming events | the source sends a change notification itself | when low-latency reaction is required and the provider documents retries |
| Periodic batch | retrieve the whole collection or increment once a day | large volumes and historical data |
| Website retrieval | read page content when no interface exists | fragile when HTML changes; assess terms of use and check more often |
Webhook and stream guarantees depend on the provider: events can be retried, delivered more than once, delayed, or out of order. Design the receiver to be idempotent and read the documentation of the particular source. Periodic reconciliation with an API or export is often an additional completeness safeguard.
Where to store it
For time-organised measurement and event data, one option is TimescaleDB, a PostgreSQL extension we use in our analytics and prediction platform. It is not required for every project: a simple volume can remain in ordinary PostgreSQL, while other workloads can justify a column store or streaming service.
Decide three things at the start:
- Keep raw and processed data separate when needed. Limit retention by purpose, cost, contract, and data-minimisation principle; do not retain a full raw record forever “just in case”.
- How long to retain detail. You can keep hourly or daily aggregates and delete older detail under an explicit rule. TimescaleDB supports retention policies and separate aggregate retention. Source: TimescaleDB retention documentation.
- How to label uncertain data. A delayed or estimated value must be recognisable.
Why collectors fail silently
Five frequent real-world causes:
- the provider changed a format or interface;
- a token or password expired;
- data scope changed, so technically valid results contain fewer records;
- a request limit was exceeded and some calls were rejected;
- local time, UTC, and daylight-saving changes created gaps or duplicates because the time model was implicit.
Format checks are not enough; you need semantic checks. Alerts can react to missing data in an expected window, an unusual record count, or a value outside an agreed range. Test thresholds so the team is not flooded with false alarms.
How we implement it
- Inventory sources — what, where, by which path, how often it changes, and which limits apply.
- Model the data — what a record is, its timestamp, and how duplicates are detected.
- Build one collector end to end, including alerts. Never five collectors without alerts.
- Add source-specific semantic checks.
- Create a status panel showing whether every source responds and when it last supplied data.
- Only then add analytics — a dashboard, reports, and forecasts.
When a collector is not needed
- Data is already in one well-structured place. A reporting tool connected to the database may be enough.
- The question is rare and needs no alert. A controlled on-demand export can cost less than a continuously maintained process.
- Nobody knows which numbers are needed yet. Collecting everything in advance creates storage and work with no decision at the end. First agree the decisions, as in vanity metrics.
- The source has no stable retrieval method. Reading a page may be the only option, but accept that it will break.
How much does it cost?
Collectors are usually part of an analytics project, not a separate purchase. Our August 2026 ranges: dashboard with several source integrations 15,000–40,000 PLN net, platform with collectors and automated reporting 40,000–120,000 PLN net, and maintenance 2,000–7,000 PLN net/month.
Pricing moves with source count, interface quality, volume, required freshness, retention, and recovery guarantees. For every decision, agree the maximum acceptable delay rather than automatically choosing real time.
Frequently asked questions
Is a ready-made data-integration tool enough?
For simple sources and small volumes, often yes. Custom collectors make sense when data needs non-standard processing, volume is large, or downtime costs money. See system integration without a programmer for the boundary.
What if the source has no API?
Check file export, read-only database access, an event queue, or an official partner integration. Website retrieval is the last option: it needs permission for that use, analysis of terms, and acceptance of greater sensitivity to HTML changes.
How long should historical data be retained?
For as long as the specific purpose, contract, or legal obligation requires — and not longer without reason. For forecasts, needed history depends on data frequency, seasonality, structural changes, and validation method; there is no universal two-cycle rule. Aggregates may be retained longer than detail if they still serve the analysis purpose.
We operate monitored collectors in our own production systems, designed for continuous work. Interruptions happen; our experience includes detecting, diagnosing, and backfilling their data. See data and analytics and our platform.

Author
Maciej Szukalski
Founder of Condictor · systems architect · research and development
He has designed and built digital products since 2014. He specialises in architecture, research, and applications with automation and intelligence layers.
See experience and working principlesHave a problem to solve?
Let’s find the right first step
Describe your situation in a few sentences. We’ll return with questions or a concrete proposal for what comes next.
