Deduplication Guide¶
DeduplicationTracker skips ad IDs already recorded by that tracker. It supports in-memory tracking for one process and persistent SQLite-backed tracking across runs. It does not deduplicate across separate database files or different tracker instances.
In-memory mode¶
State lives in a Python set. Fast, but lost when the process exits.
from meta_ads_collector import MetaAdsCollector, DeduplicationTracker
tracker = DeduplicationTracker(mode="memory")
with MetaAdsCollector() as collector:
for ad in collector.search(query="test", dedup_tracker=tracker):
print(ad.id) # This ID is recorded before the ad is yielded.
print(f"Unique ads: {tracker.count()}")
Persistent mode¶
State is stored in a SQLite database file. Survives across process restarts.
from meta_ads_collector import MetaAdsCollector, DeduplicationTracker
tracker = DeduplicationTracker(mode="persistent", db_path="collection_state.db")
with MetaAdsCollector() as collector:
for ad in collector.search(query="test", dedup_tracker=tracker):
print(ad.id) # Skips ads already recorded in any previous run
tracker.close()
# The search generator saves yielded ad IDs when it finishes or is closed. It advances the last-collection time only when the search reaches the end without parse errors or an unmet result limit. Closing early or stopping at a limit preserves yielded IDs but leaves the last-collection time unchanged.
The SQLite database contains two tables:
- seen_ads -- maps ad IDs to their first-seen timestamp
- collection_runs -- records the timestamp of each completed collection run
Incremental collection¶
Combine persistent deduplication with date filtering to only collect new ads since the last run.
from meta_ads_collector import MetaAdsCollector, DeduplicationTracker, FilterConfig
tracker = DeduplicationTracker(mode="persistent", db_path="state.db")
last_run = tracker.get_last_collection_time()
# Only fetch ads newer than the last collection
filters = FilterConfig(start_date=last_run) if last_run else None
with MetaAdsCollector() as collector:
for ad in collector.search(query="crypto", filter_config=filters, dedup_tracker=tracker):
print(ad.id) # Replace with your application's processing code.
tracker.close()
CLI usage¶
# In-memory deduplication (within a single run)
meta-ads-collector -q "test" --dedup -o ads.json
# Persistent deduplication (across runs)
meta-ads-collector -q "test" --state-file state.db -o ads.json
# Incremental: only collect ads since the last run
meta-ads-collector -q "test" --state-file state.db --since-last-run -o new_ads.jsonl
API reference¶
DeduplicationTracker¶
from meta_ads_collector import DeduplicationTracker
memory_tracker = DeduplicationTracker(mode="memory")
persistent_tracker = DeduplicationTracker(mode="persistent", db_path="state.db")
| Method | Description |
|---|---|
has_seen(ad_id) |
Returns True if the ad ID was previously recorded |
mark_seen(ad_id, timestamp=None) |
Record an ad ID as seen |
get_last_collection_time() |
Returns the datetime of the most recent completed run, or None |
update_collection_time() |
Record the current time as the latest collection run |
save() |
Persist changes to disk (persistent mode only; no-op for memory) |
load() |
Load state from disk (persistent mode only; no-op for memory) |
count() |
Number of unique ad IDs tracked |
clear() |
Remove all tracked state |
close() |
Close the database connection (persistent mode) |
Context manager¶
DeduplicationTracker supports with for automatic save and close:
from meta_ads_collector import DeduplicationTracker
with DeduplicationTracker(mode="persistent", db_path="state.db") as tracker:
tracker.mark_seen("12345")
# Automatically saves and closes on exit