Repository navigation
Use takeWhile method from Range - #720
EnverOsmanov wants to merge 1 commit into
Conversation
|
Hi @EnverOsmanov, thanks for the PR! val lastCellNum = r.getLastCellNum
colInd
.iterator
.filter(_ < lastCellNum) |
|
If Benchmarks: Here is the code how I read the data. Btw, I just checked the content of |
|
The alternative approach to avoid iteration over full But I'm not exactly sure what was the idea behind the change in V2. |
|
Hmm, maybe it is to be able to do the r.getCell(_, MissingCellPolicy.CREATE_NULL_AS_BLANK)@quanghgx could you chime in here? |
|
If |
6b58ec4 to
6866cb1
Compare
When dataAddress specifies only a starting cell, colInd becomes a huge range (e.g. 1 to 16383). The previous code iterated the entire range per row using .filter(), causing O(rows * maxColumns) comparisons. Replace with direct range clipping: colInd.start to min(colInd.last, lastCellNum - 1). This is O(1) per row and also hoists the getLastCellNum call out of the per-element evaluation. Fixes #720 Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
Superseded by #1032 which uses direct range clipping instead of |
* perf: Clip column range instead of filtering in V2 DataLocator When dataAddress specifies only a starting cell, colInd becomes a huge range (e.g. 1 to 16383). The previous code iterated the entire range per row using .filter(), causing O(rows * maxColumns) comparisons. Replace with direct range clipping: colInd.start to min(colInd.last, lastCellNum - 1). This is O(1) per row and also hoists the getLastCellNum call out of the per-element evaluation. Fixes #720 Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Update src/main/scala/dev/mauch/spark/excel/v2/DataLocator.scala Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
The symptoms:
I have a file with ~1 million rows, 125 columns. It takes ~12 seconds to count lines with spark-excel's API V1 and ~2 minutes with API V2.
The issue:
Rangedoes not contain own optimized methodfilter, that's why it uses method fromTraversableLikewhich iterates over each number in range.r.getLastCellNumevaluated for each number in range.Here are some rough benchmarks with another file:
filter => 50 seconds
val lastCellNum => 38 seconds
withFilter => 20 seconds
takeWhile => 12 seconds
API V1 => 12 seconds
(File taken from here and manually converted to "xlsx")
PS. API V2 seems great! :)