Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
150 commits
Select commit Hold shift + click to select a range
6461c6e
updated the blog
nayanj98 Mar 16, 2026
a79edd8
feat: optimization documentation (#393)
hash-data Apr 8, 2026
b1d63df
Merge pull request #385 from datazip-inc/doc/optimization-fusion-config
saptarshi-datazip Apr 9, 2026
eac6abe
release notes updated with v0.6.1 to 0.6.4 (#394)
nayanj98 Apr 10, 2026
c13e8a0
docs:Update data filter section with CDC filter (#373)
nayanj98 Apr 10, 2026
4a9bf63
chore: add doc for cli testing in olake-ui (#351)
vishalm0509 Apr 14, 2026
269143c
fix: sidebar issues (#401)
deepanshupal09-datazip Apr 15, 2026
657e219
update mysql data types (#398)
nayanj98 Apr 20, 2026
5a72e92
docs: Add API documentation (#402)
nayanj98 Apr 20, 2026
35d41b1
docs: Updated destination docs with Arrow Writer (#397)
nayanj98 Apr 20, 2026
dcbaf74
docs: Updated postgres source setup (#395)
nayanj98 Apr 20, 2026
4783673
docs: Update job scheduling using cron (#399)
nayanj98 Apr 24, 2026
0152e6f
docs: Update release notes with Optimzation section (#400)
nayanj98 Apr 24, 2026
7254ad4
docs:Bulk configuration section added in Ingestion Properties (#407)
nayanj98 Apr 27, 2026
f3d10d7
blog: OLake Fusion vs Spark Compaction (#404)
nayanj98 Apr 28, 2026
d3047f6
docs: Optimization Benchmarking Document (#406)
nayanj98 Apr 28, 2026
76d94d1
docs: Update Creating your First Job Pipeline (#396)
nayanj98 Apr 29, 2026
1e7fd54
improvement: renamed optimization to either maintenance or compaction…
siddharth-chevella Apr 30, 2026
28517af
docs: refactor compose quick start and upgrade tabs (#410)
saptarshi-datazip May 2, 2026
4ab15e6
updated fusion setup links (#411)
siddharth-chevella May 4, 2026
43b5bb8
title-update (#412)
siddharth-chevella May 5, 2026
864197a
update release notes for olake go and fusion (#413)
nayanj98 May 5, 2026
d160045
Remove normalisation false mandatory info from data partitioning (#415)
nayanj98 May 6, 2026
94780b6
docs: Added Schema parameter to postgres docs (#416)
nayanj98 May 8, 2026
51ae9a7
update release notes for ingestion with v0.7.1 (#417)
nayanj98 May 8, 2026
e71c57a
update go release notes with v0.7.2 (#420)
nayanj98 May 12, 2026
bc649a5
OLake Fusion SEO Optimizations (#419)
anshika-oss May 13, 2026
8138800
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 13, 2026
0a81ed0
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 13, 2026
b3cefaa
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 13, 2026
0ee4ea8
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 13, 2026
37c8fb8
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 13, 2026
8f77570
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 13, 2026
ccf5b13
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 13, 2026
3e18900
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 14, 2026
2e22a23
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 14, 2026
24898fd
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 14, 2026
9bd96d6
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 14, 2026
234fa2b
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 14, 2026
42b4681
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 14, 2026
4045325
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 15, 2026
700586d
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 15, 2026
aa067db
Add blog: Conflict-Free CDC into Apache Iceberg
shuvajyotikar13 May 15, 2026
f1fc01b
Merge pull request #421 from shuvajyotikar13/master
nayanj98 May 15, 2026
4b303bd
doc: fix helm installation doc user flow (#418)
saptarshi-datazip May 20, 2026
c2b24d0
added go v0.7.3 (#423)
nayanj98 May 20, 2026
bc2e68d
docs: Update postgres docs with info to retain full row (#422)
nayanj98 May 20, 2026
9b9680b
docs: Updated the Overview page for Ingestion (OLake Go) (#414)
nayanj98 May 20, 2026
f4a9558
docs: Update MSSQL doc with read only replica (#424)
nayanj98 May 22, 2026
95eea49
Merge latest master into blog/debeziumvsolake_update
nayanj98 May 22, 2026
cb2d849
updated content based on review
nayanj98 May 22, 2026
4aeac87
chore: add olake fusion announcement banner (#425)
deepanshupal09-datazip May 25, 2026
8396667
docs: Use Apache Amoro™ branding in blog content (#426)
nayanj98 May 27, 2026
3399a04
docs: refactored minor fusion helm doc portion (#427)
saptarshi-datazip May 28, 2026
07ceeda
docs: Updated release notes with v0.7.4 and mysql benchmark (#428)
nayanj98 Jun 2, 2026
b94edf0
update data type for postgres source
nayanj98 Jun 2, 2026
a5c0669
Merge pull request #429 from datazip-inc/docs/postgres_source_update
nayanj98 Jun 2, 2026
94136b2
doc: add section for external temporal setup (#431)
schitizsharma Jun 5, 2026
af06d4e
docs: MSSQL benchmarks update (#432)
nayanj98 Jun 11, 2026
7b6eb94
Add blog: Apache Iceberg Row Lineage: Tracking Data Lineage at the Ro…
anshika-oss Jun 13, 2026
5978892
docs: Restructure for OLake Go and Fusion (#430)
nayanj98 Jun 14, 2026
f609d3c
Add Xeno customer story: AWS DMS alternative for MySQL CDC
anshika-oss Jun 14, 2026
ed72a1a
Address PR review comments on Xeno customer story
anshika-oss Jun 15, 2026
8153f32
chore: cleanup testinmonialcard code
deepanshupal09-datazip Jun 15, 2026
b8a0fc3
Merge pull request #435 from datazip-inc/customer-stories/xeno-aws-dm…
anshika-oss Jun 15, 2026
1b063a8
update mssql documentation
nayanj98 Jun 16, 2026
63c10be
blog: add AWS DMS vs OLake comparison post
anshika-oss Jun 16, 2026
7598aa8
Merge pull request #437 from datazip-inc/blog/aws-dms-vs-olake
anshika-oss Jun 17, 2026
c087057
context correction
nayanj98 Jun 17, 2026
f91ddc0
Update release notes for Go and Fusion
nayanj98 Jun 17, 2026
62fc58a
Merge pull request #438 from datazip-inc/docs/release_go_0.7.6_fusion…
hash-data Jun 18, 2026
7345075
update troubleshooting
nayanj98 Jun 18, 2026
b58ec29
Merge branch 'master' into docs/mssql_update
nayanj98 Jun 18, 2026
873526f
fusion bulk edit documentation
nayanj98 Jun 18, 2026
3a18392
Update the manage instance documentation while using secondaty database
nayanj98 Jun 19, 2026
34eae6d
update the ssh tunnel note
nayanj98 Jun 19, 2026
d846088
change heading for single table congifuration
nayanj98 Jun 19, 2026
e77eef7
update definition for olake imported catalogs
nayanj98 Jun 19, 2026
f9d708f
Merge branch 'master' into blog/debeziumvsolake_update
nayanj98 Jun 19, 2026
920619e
doc: add documentation for k8s
schitizsharma Jun 19, 2026
3b02d2b
doc: add container info
schitizsharma Jun 19, 2026
efb7a2d
chore: add steps for docker login
schitizsharma Jun 19, 2026
bda6497
Merge pull request #436 from datazip-inc/docs/mssql_update
nayanj98 Jun 19, 2026
da52fba
chore: add private registry to docker compose
schitizsharma Jun 19, 2026
dcc433c
docs: Update docs with source name convention (#441)
nayanj98 Jun 19, 2026
ceac312
chore: re placing of block
schitizsharma Jun 19, 2026
9df95d3
Merge pull request #442 from datazip-inc/doc/k8s-private-registry-not…
schitizsharma Jun 19, 2026
f39468c
Merge pull request #439 from datazip-inc/docs/fusion_bulk_edit
nayanj98 Jun 22, 2026
e391c5e
update postgres troubleshooting
nayanj98 Jun 22, 2026
94d0a4a
change troubleshooting title
nayanj98 Jun 22, 2026
4be7698
Merge pull request #443 from datazip-inc/docs/read_replica_update_pos…
nayanj98 Jun 22, 2026
daadf23
kafka document update
nayanj98 Jun 22, 2026
4464a09
add note to redirect to fusion in go
nayanj98 Jun 23, 2026
33c228e
Merge pull request #382 from datazip-inc/blog/debeziumvsolake_update
nayanj98 Jun 23, 2026
2bc22e8
link change for compaction first pipeline
nayanj98 Jun 23, 2026
5bb061c
update the note for consumer group
nayanj98 Jun 23, 2026
28b46bc
Merge pull request #444 from datazip-inc/docs/kafka_doc_update
nayanj98 Jun 23, 2026
bcfdf14
docs: Update the slack links in the documentation (#446)
nayanj98 Jun 25, 2026
2f55eea
FAQs for all the blogs
anshika-oss Apr 7, 2026
df80074
doc: add pod annotation doc and rectify the image pull secret jobSA (…
schitizsharma Jun 25, 2026
a20a98e
update release notes OLake Go v0.7.7
nayanj98 Jun 26, 2026
adcfcb3
update source naming convention point
nayanj98 Jun 29, 2026
a3e10f2
mention cli and ui specifically
nayanj98 Jun 29, 2026
fe25aba
chore: add setting for fusion db (#449)
schitizsharma Jun 29, 2026
ea6b022
add go v0.7.8
nayanj98 Jun 30, 2026
fe7c0e4
Merge pull request #447 from datazip-inc/docs/go_v0.7.7_release_note
nayanj98 Jun 30, 2026
af087e0
docs: Update clear destination command (#448)
nayanj98 Jul 2, 2026
29eccdb
chore: document S3 rolling param
mihir-datazip Jul 2, 2026
385b9d5
update fusion release notes (#452)
nayanj98 Jul 6, 2026
76f1b3d
release notes go v0.8.0
nayanj98 Jul 8, 2026
e6c3f44
Address review comments across blog FAQs
anshika-oss Jul 8, 2026
712e1e4
add release notes go_v0.8.1
nayanj98 Jul 9, 2026
8e17400
Merge pull request #454 from datazip-inc/docs/release_go_v0.8.0
nayanj98 Jul 9, 2026
a0f9cb7
Address review comments across blog FAQs
anshika-oss Jul 9, 2026
c8fbc74
Fix illegible company name text on customer story cards
anshika-oss Jul 9, 2026
1470343
Merge pull request #456 from datazip-inc/fix/customer-card-company-na…
anshika-oss Jul 9, 2026
9f65f46
Update static/llm.txt
anshika-oss Jul 13, 2026
4e79a36
update release notes go
nayanj98 Jul 14, 2026
c04ca5e
Merge pull request #460 from datazip-inc/docs/release_go_0.8.2
nayanj98 Jul 14, 2026
f467e77
doc: add fusion compaction scheduling logic doc (#461)
schitizsharma Jul 14, 2026
20dc4da
blog: one time delivery blog for iceberg (#458)
vaibhav-datazip Jul 15, 2026
a967cac
Add Google Analytics tag and fix broken headTags config
anshika-oss Jul 15, 2026
8486826
Second review comments resolved, made sure all the FAQs are titled FA…
anshika-oss Jul 20, 2026
86604b9
Removed em-dashes and fixed the minor comments, also made sure the pa…
anshika-oss Jul 21, 2026
2a0dca1
moved all FAQs to the bottom before the outro line
anshika-oss Jul 22, 2026
b651848
Merge pull request #391 from anshika-oss/FAQ-all-blogs
anshika-oss Jul 22, 2026
27bdf19
Add blog post: How OLake Handles Schema Evolution (#464)
anshika-oss Jul 23, 2026
e34aba5
Update llm.txt: fix connector links, benchmark entries, meetup descri…
anshika-oss Jul 24, 2026
e93af7a
Merge pull request #459 from datazip-inc/update-llm-txt
anshika-oss Jul 24, 2026
e1831cd
Use plugin-google-gtag preset option and trim duplicate SEO headTags
anshika-oss Jul 27, 2026
527863d
added the twitter handle
anshika-oss Jul 28, 2026
34e1abc
Merge pull request #463 from datazip-inc/feat/google-analytics-manual
anshika-oss Jul 28, 2026
5173574
docs: Query engine compatibility doc update (#451)
nayanj98 Jul 29, 2026
84d0d8c
update release notes olake go
nayanj98 Jul 30, 2026
6678f2c
Merge pull request #468 from datazip-inc/docs/release_go_v0.9.0
nayanj98 Jul 30, 2026
7190270
update rest with biglake catalog
nayanj98 Jul 30, 2026
1f04449
feat(docs): bounded, sticky tab bars with nested tab stacking (#465)
shubham19may Jul 30, 2026
668546d
add the version requirement
nayanj98 Jul 31, 2026
a7651e1
add proper description for table
nayanj98 Jul 31, 2026
b94af23
update the table and remove emoji from unity heading
nayanj98 Jul 31, 2026
99495f9
update table with s3 path
nayanj98 Aug 3, 2026
8e7f8f3
Merge pull request #469 from datazip-inc/docs/biglake_catalog
nayanj98 Aug 3, 2026
8e72cbf
fix the sidebar for biglake
nayanj98 Aug 3, 2026
58fd618
change destination.json
nayanj98 Aug 3, 2026
a5facc9
Merge pull request #471 from datazip-inc/docs/update_biglake_catalog
nayanj98 Aug 3, 2026
7f99266
feat: new home landing page, fusion page and go page (#472)
deepanshupal09-datazip Aug 5, 2026
4dfc3f7
chore: address comments
mihir-datazip Aug 5, 2026
35d3c71
chore: clarify fractional max file size in CLI docs
mihir-datazip Aug 6, 2026
fad49b4
Merge pull request #450 from mihir-datazip/s3-rolling-config
nayanj98 Aug 7, 2026
e26c449
docs(mongodb): document inline TLS certificate fields for UI/Docker
krishanu7 Aug 9, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
Original file line number Diff line number Diff line change
Expand Up @@ -336,48 +336,67 @@ Are you interested in unlocking the full potential of your data without the need
With features like data ingestion from 150+ sources including MongoDB connectors, data warehousing, data analytics, and data transformation solutions, Datazip can help you make fast, data-driven decisions.

## FAQs
### Q1. What are the most common MongoDB ETL errors and how do you diagnose them?
- **Connection timeout errors** — Check network connectivity, firewall/security group rules blocking port 27017, and MongoDB authentication credentials
- **Schema validation failures** — Caused by polymorphic fields or missing required fields across documents in the same collection
- **Data type mismatch errors** — Where the source field type differs from the target column type
- **Socket timeout (`socketTimeoutMS`) exhaustion during large collection scans** — Occurs when MongoDB takes longer than the configured `socketTimeoutMS` to respond to a query, common during unoptimized aggregate queries or large full-collection reads. Increase `socketTimeoutMS` in your connection settings and ensure queries are properly indexed to avoid full collection scans.

### Q2. How should I set up MongoDB for ETL to minimize pipeline errors?
Best practices for an ETL-ready MongoDB setup include:

- **Enable Read Preference on secondary nodes** to offload ETL reads from the primary and avoid impacting operational performance
- **Create indexes on user-defined timestamp fields** (such as an application-managed `updated_at` field) that are used for cursor-based incremental sync — note this is not a built-in MongoDB field and must be maintained by your application
- **Set `socketTimeoutMS` and `serverSelectionTimeoutMS`** appropriately per operation for long-running collection reads, keeping in mind these are per-operation settings, not session-level configurations in most drivers
- **Configure oplog retention** to cover at least 24–48 hours of changes to ensure CDC consumers do not fall behind the retention window
- **Ensure a replica set is configured:** this is a hard requirement for change streams and oplog-based CDC; standalone MongoDB instances do not support these features

### Q3. What causes connection timeout errors in MongoDB ETL pipelines and how do I fix them?
Connection timeouts typically occur due to:

- **Network/firewall issues:** Firewall or security group rules blocking the ETL tool's IP from reaching MongoDB on port 27017
- **Authentication failures:** Wrong credentials, incorrect `authSource` database, or the user lacking required permissions
- **Connection pool exhaustion:** Too many concurrent ETL workers exceeding the `maxPoolSize` setting, or connection leaks in application code causing "server selection timed out" errors
- **SSL/TLS configuration mismatches:** The ETL tool lacking the correct CA certificate to validate the MongoDB server's SSL certificate

**Recommended debug approach:** Test connectivity directly with the MongoDB shell (`mongosh`) using the same connection string first. If that succeeds, the issue is in your ETL tool's configuration — verify credentials, SSL settings, and connection string parameters. If the shell also fails, the issue is at the network or DNS level.

### Q4. How do I handle schema validation errors when MongoDB documents have inconsistent structures?
Schema validation errors occur because MongoDB allows polymorphic data — documents with varying structures or different data types for the same field — within a single collection. Solutions include:

- **Use schema inference with adequate sampling** — Increase the sample size when inferring the schema so the ETL tool captures the full range of field variations, rather than relying on a small, potentially unrepresentative subset
- **Mark fields as nullable/optional** for fields that may be absent in some documents
- **Apply type coercion rules** to handle polymorphic fields by enforcing a consistent target type during ingestion
- **Filter or quarantine malformed documents** using pre-ingestion validation rules — MongoDB also supports `validationAction: "warn"` mode, which logs invalid documents without rejecting them, making it a useful diagnostic tool during ETL pipeline development
- **Use a compatible ETL tool** that natively supports MongoDB's BSON types (including `Decimal128`, `ObjectID`) and flexible schema evolution

### Q5. What are best practices for MongoDB ETL setup in production environments?
For production MongoDB ETL pipelines:

- **Use a dedicated read-only ETL user** with the minimum permissions required — typically `read` on source collections and `clusterMonitor` for oplog access
- **Connect to a replica set secondary** to avoid adding read load to the primary node
- **Implement checkpointing using resume tokens** so failed syncs resume from the last successfully processed oplog position rather than restarting from scratch — store the resume token durably and pass it back on reconnection
- **Monitor oplog lag actively** — a small oplog (e.g., 1GB on a high-throughput cluster) may only retain a few hours of changes; if your CDC consumer falls behind the retention window, you will need to trigger a full resync
- **Test oplog partial-update handling in staging** before deploying to production — MongoDB's `$set` update operator produces partial update events in the oplog (not full document replacements), and many ETL tools handle these differently; validate that your tool correctly reconstructs the full document from partial oplog events before going live

<Faq showHeading={false} data={[
{
question: "Q1. What are the most common MongoDB ETL errors and how do you diagnose them?",
answer: <ul>
<li><strong>Connection timeout errors.</strong> The job fails before reading any data. Diagnose by checking network connectivity to the host, firewall or security group rules blocking port 27017, and your MongoDB authentication credentials.</li>
<li><strong>Schema validation failures.</strong> Records get rejected at the destination. Diagnose by sampling documents in the collection and looking for polymorphic fields or required fields missing across documents.</li>
<li><strong>Data type mismatch errors.</strong> A column fails to write or values get coerced unexpectedly. Diagnose by comparing the source field type against the target column type for the field named in the error.</li>
<li><strong>Socket timeout (<code>socketTimeoutMS</code>) exhaustion during large collection scans.</strong> A query runs longer than the configured <code>socketTimeoutMS</code> and the connection drops mid-scan. Diagnose by comparing query duration against <code>socketTimeoutMS</code> and checking whether the query uses an index. Raise <code>socketTimeoutMS</code> and add indexes so reads avoid full collection scans.</li>
</ul>
},
{
question: "Q2. How should I set up MongoDB for ETL to minimize pipeline errors?",
answer: <div>
<p>Best practices for an ETL-ready MongoDB setup include:</p>
<ul>
<li><strong>Enable Read Preference on secondary nodes</strong> to offload ETL reads from the primary and avoid impacting operational performance</li>
<li><strong>Create indexes on user-defined timestamp fields</strong> (such as an application-managed <code>updated_at</code> field) used for cursor-based incremental sync</li>
<li><strong>Set <code>socketTimeoutMS</code> and <code>serverSelectionTimeoutMS</code></strong> appropriately per operation for long-running collection reads</li>
<li><strong>Configure oplog retention</strong> to cover at least 24–48 hours of changes to ensure CDC consumers do not fall behind the retention window</li>
<li><strong>Ensure a replica set is configured</strong> this is a hard requirement for change streams and oplog-based CDC; standalone MongoDB instances do not support these features</li>
</ul>
</div>
},
{
question: "Q3. What causes connection timeout errors in MongoDB ETL pipelines and how do I fix them?",
answer: <div>
<p>Connection timeouts typically occur due to:</p>
<ul>
<li><strong>Network/firewall issues.</strong> Firewall or security group rules blocking the ETL tool's IP from reaching MongoDB on port 27017. Fix: allowlist the ETL tool's IP or subnet on port 27017 and confirm the route with <code>telnet host 27017</code> or <code>nc -zv host 27017</code>.</li>
<li><strong>Authentication failures.</strong> Wrong credentials, incorrect <code>authSource</code> database, or the user lacking required permissions. Fix: verify the credentials, set <code>authSource</code> to the database where the user was created (often <code>admin</code>), and grant the user read access to the source collections.</li>
<li><strong>Connection pool exhaustion.</strong> Too many concurrent ETL workers exceeding the <code>maxPoolSize</code> setting, or connection leaks causing "server selection timed out" errors. Fix: raise <code>maxPoolSize</code> to cover your worker count, or lower worker concurrency so it stays within the pool, and make sure connections are closed after each task.</li>
<li><strong>SSL/TLS configuration mismatches.</strong> The ETL tool lacking the correct CA certificate to validate the MongoDB server's SSL certificate. Fix: point the connection at the right CA bundle with <code>tlsCAFile</code>, and confirm the host matches the certificate's subject so validation passes.</li>
</ul>
<p><strong>Recommended debug approach:</strong> Test connectivity directly with the MongoDB shell (<code>mongosh</code>) using the same connection string first. If that succeeds, the issue is in your ETL tool's configuration. If the shell also fails, the issue is at the network or DNS level.</p>
</div>
},
{
question: "Q4. How do I handle schema validation errors when MongoDB documents have inconsistent structures?",
answer: <div>
<p>Schema validation errors occur because MongoDB allows polymorphic data within a single collection. Solutions include:</p>
<ul>
<li><strong>Use schema inference with adequate sampling.</strong> Increase the sample size so the ETL tool captures the full range of field variations</li>
<li><strong>Mark fields as nullable/optional</strong> for fields that may be absent in some documents</li>
<li><strong>Apply type coercion rules</strong> to handle polymorphic fields by enforcing a consistent target type during ingestion</li>
<li><strong>Filter or quarantine malformed documents</strong> using pre-ingestion validation rules. MongoDB's <code>validationAction: "warn"</code> mode logs invalid documents without rejecting them, useful during pipeline development</li>
<li><strong>Use a compatible ETL tool</strong> that natively supports MongoDB's BSON types (including <code>Decimal128</code>, <code>ObjectID</code>) and flexible schema evolution</li>
</ul>
</div>
},
{
question: "Q5. What are best practices for MongoDB ETL setup in production environments?",
answer: <div>
<p>For production MongoDB ETL pipelines:</p>
<ul>
<li><strong>Use a dedicated read-only ETL user</strong> with minimum permissions required, typically <code>read</code> on source collections and <code>clusterMonitor</code> for oplog access</li>
<li><strong>Implement checkpointing using resume tokens</strong> so failed syncs resume from the last successfully processed oplog position. Store the resume token durably and pass it back on reconnection</li>
<li><strong>Monitor oplog lag actively.</strong> A small oplog (e.g., 1GB on a high-throughput cluster) may only retain a few hours of changes; if your CDC consumer falls behind, you will need to trigger a full resync</li>
<li><strong>Test oplog partial-update handling in staging</strong> before deploying. MongoDB's <code>$set</code> operator produces partial update events (not full document replacements), and many ETL tools handle these differently</li>
</ul>
</div>
},
]} />

<BlogCTA/>
65 changes: 64 additions & 1 deletion blog/2024-09-16-mongodb-etl-challenges.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -563,7 +563,70 @@ By following best practices, such as using CDC, batching, and data validation, c

*4. MongoDB Documentation, "Working with Nested Data," MongoDB.*


## FAQs
<Faq showHeading={false} data={[
{
question: "Q1. What are the main MongoDB ETL challenges when moving data to a data warehouse?",
answer: <div>
<p>The four key challenges are:</p>
<ol>
<li><strong>Schema flexibility:</strong> MongoDB's schema-less design creates inconsistent field structures that clash with rigid warehouse schemas</li>
<li><strong>Large initial loads:</strong> Must be parallelized and checkpointed to handle terabyte-scale collections reliably</li>
<li><strong>Changing data types (polymorphic keys):</strong> The same field can appear as different types across documents (e.g., <code>age</code> as an integer in one document and a string in another)</li>
<li><strong>Complex nested fields and arrays:</strong> Must be transformed into a flat relational format without causing row explosion or data duplication</li>
</ol>
</div>
},
{
question: "Q2. How does MongoDB's schema flexibility create problems during ETL to structured systems?",
answer: <div>
<p>MongoDB allows any document to omit fields or use different types for the same field across documents. When moving this data to relational warehouses or Iceberg tables that require consistent schemas, you encounter:</p>
<ul>
<li><strong>Type mismatches:</strong> A field that is an integer in some documents and a string in others</li>
<li><strong>Missing values:</strong> Sparse fields that exist in only a subset of documents require NULL-filling across the rest</li>
<li><strong>Inconsistent nesting structures:</strong> The same logical field may appear as a simple string in one document and a nested object in another</li>
</ul>
<p>ETL pipelines must detect and resolve these variations without silently dropping or corrupting data.</p>
</div>
},
{
question: "Q3. What is the best approach for handling the first full load of a large MongoDB collection?",
answer: <div>
<p>For large collections, parallelize the initial load using <strong><code>_id</code>-based range queries or bucket-based partitioning</strong> to split the collection into independent read ranges, then load those ranges concurrently across multiple worker threads.</p>
<p>Key practices to follow:</p>
<ul>
<li><strong>Implement checkpointing</strong> so that if the load fails mid-way, it resumes from the last successfully completed chunk rather than restarting from scratch</li>
<li><strong>Read from a replica set secondary</strong> to avoid load on the primary during the bulk read</li>
<li><strong>Record the oplog timestamp or resume token</strong> at the start of the full load so that incremental CDC replication can pick up exactly from that point once the snapshot completes</li>
</ul>
<p>After the full load completes, switch to <strong>change streams or oplog-based CDC</strong> for ongoing incremental replication.</p>
</div>
},
{
question: "Q4. How should you handle array fields when doing MongoDB ETL to a relational target?",
answer: <div>
<p>Arrays in MongoDB documents should generally be <strong>exploded into separate child tables</strong> with foreign key references to the parent document. Strategies by array type:</p>
<ul>
<li><strong>Arrays of simple values:</strong> Create a child table with a <code>parent_id</code> column and a <code>value</code> column. Each array element becomes one row.</li>
<li><strong>Arrays of complex objects:</strong> Each object in the array becomes a full row in the child table, with all object fields mapped to columns and a foreign key back to the parent.</li>
</ul>
<p><strong>Avoid flattening arrays inline:</strong> this causes row explosion and massive data duplication, making the final dataset several times larger than the original.</p>
<p>For arrays you do not need for analytics, consider skipping them entirely during extraction rather than flattening unnecessarily.</p>
</div>
},
{
question: "Q5. How do you handle polymorphic data types in MongoDB ETL pipelines?",
answer: <div>
<p>Polymorphic fields, where the same key holds values of different types across documents, can be addressed using one of these strategies:</p>
<ul>
<li><strong>Type promotion.</strong> Promote all values to the most permissive compatible type (e.g., convert both integers and strings to <code>string</code>). Use numeric type promotions where safe (e.g., <code>int</code> to <code>long</code>, <code>float</code> to <code>double</code>). Best when the types are compatible and you want a single clean column. Avoid it when promoting to <code>string</code> would lose type information you need downstream.</li>
<li><strong>Separate typed columns.</strong> Create distinct columns per data type (e.g., <code>age_int</code> and <code>age_string</code>). Older data stays in the original column; new data with a different type populates the new column. Best when you need to preserve every value losslessly and downstream consumers can handle querying across columns. The tradeoff is a wider, sparser table.</li>
<li><strong>Schema inference with sampling.</strong> Run a sampling step across the collection before defining your pipeline schema, to determine the dominant type for each field and surface polymorphic fields early. Best as a first step before any pipeline run, so you choose the right handling strategy instead of discovering type conflicts mid-sync.</li>
<li><strong>JSON/variant column.</strong> Store the field as a semi-structured column and handle parsing in downstream transformations. Best when types are unpredictable or change often and you want to defer parsing. Only available when your target warehouse natively supports it, such as Snowflake's <code>VARIANT</code>, BigQuery's <code>JSON</code>, or Redshift's <code>SUPER</code>.</li>
</ul>
</div>
},
]} />

I’d love to hear your thoughts about this, so feel free to reach out to me on [LinkedIn](https://www.linkedin.com/in/zriyansh/).

Expand Down
Loading