Skip to content

fetch data from leader - #770

Closed
JacksonYao287 wants to merge 1 commit into
eBay:masterfrom
JacksonYao287:retry-fetch-data
Closed

JacksonYao287 wants to merge 1 commit into
eBay:masterfrom
JacksonYao287:retry-fetch-data

Conversation

@JacksonYao287

Copy link
Copy Markdown
Member

1 fetch data from leader , not leader, to avoid the case that originator is the out_member when replace member happens

2 if error happens when handling fetch_data request, return empty mesage and let follower retry. don`t crash

@codecov-commenter

codecov-commenter commented Jul 14, 2025 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 33.33333% with 12 lines in your changes missing coverage. Please review.
✅ Project coverage is 66.72%. Comparing base (1a0cef8) to head (5218d1c).
⚠️ Report is 274 commits behind head on master.

Files with missing lines Patch % Lines
src/lib/replication/repl_dev/raft_repl_dev.cpp 25.00% 10 Missing and 2 partials ⚠️
❗ Your organization needs to install the Codecov GitHub app to enable full functionality.
Additional details and impacted files
@@             Coverage Diff             @@
##           master     #770       +/-   ##
===========================================
+ Coverage   56.51%   66.72%   +10.20%     
===========================================
  Files         108      110        +2     
  Lines       10300    13056     +2756     
  Branches     1402     1895      +493     
===========================================
+ Hits         5821     8711     +2890     
+ Misses       3894     3347      -547     
- Partials      585      998      +413     

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@JacksonYao287
JacksonYao287 requested a review from Besroy July 14, 2025 13:38

@xiaoxichen xiaoxichen left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

It will be interesting if we allow to fetch from non-leader which will offload the IO pressure of fetch_data on leader. cc @Besroy

But that can be done in other track.

RD_REL_ASSERT(false, "Error in reading data");

// if read data failed, we should ignore the rpc_data and let the follower retry the fetch
RD_LOGT(NO_TRACE_ID,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lets have a LOGD?

return;

// TODO: Find a way to return error to the Listener
// TODO: actually will never arrive here as iomgr will assert

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove the TODO and put the non-io-error cases into the L1305

Comment thread src/lib/replication/repl_dev/raft_repl_dev.cpp Outdated
group_msg_service()
->data_service_request_bidirectional(
originator, FETCH_DATA,
leader_server_id, FETCH_DATA,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove the comment in L1207

rreq->remote_blkid().server_id /* blkid_originator */,
leader_server_id,
builder->CreateVector(rreq->remote_blkid().blkid.serialize().cbytes(),
rreq->remote_blkid().blkid.serialized_size())));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the originator != leader, will the remote_blkid change (due to garbage and vchunk->pchunk mapping)?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

HO will read from the given remote blkid and verify data, if failed, it will try to read from index.
This fallback allow us send fetch data to anyone.

@Besroy Besroy Jul 16, 2025 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your explanation. I'm not sure if we need to add some logic in HO to handle scenarios like chunk reclamation or disk bad here. Nevertheless, it looks good to me.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chunk reclaim is fine as validate_blob will fail. Bad drive (or degrade mode) is a concern..... lets create an issue and wait zhiteng's POC on single-disk-mode....

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created a issue eBay/HomeObject#330 in HO

COUNTER_INCREMENT(m_metrics, fetch_total_blk_size, total_size);
if (!raw_data || total_size == 0) {
RD_LOGW(NO_TRACE_ID, "Data Channel: FetchData returned empty payload, ignoring");
return;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

JFYI: If a fetch is called during the end_of_append_batch and the leader returns empty data, the follower currently does not retry. After expiration, it will trigger the assertion at notify_after_data_written:HS_REL_ASSERT(false, "Data fetch timeout, should not happen");. However, this seems to be an existing issue; perhaps we could add a TODO for a retry mechanism.

@JacksonYao287

Copy link
Copy Markdown
Member Author

@Besroy @xiaoxichen

I have changed the policy of fetching data to fetch data from a random peer. as a result , it need the host to be capable of parsing user_header and get the correct blob. However, In the current raft_repl_dev UT, it use the default implementation of fetch_data, which will directly fetch data by blk_id and not do any user_header parsing, so this will lead to some data verification error. to support this case, I also add a choice to also support the case of only fetching data from originator.

@xiaoxichen xiaoxichen left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm aside from a nit

// clear reqs that has allocated blks on the given chunk.
void clear_chunk_req(chunk_num_t chunk_id);

static void enable_fetch_data_only_from_originator(bool enable) { m_fetch_data_only_from_originator = enable; }

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggest add virtual bool ReplDevListener::support_fetch_from_non_leader() {return false}

It will be cleaner.
In HO we overwrite this member function to return true.

std::vector< int32_t > peer_ids;
for (const auto& srv_config : srv_configs) {
auto peer_id = srv_config->get_id();
if (peer_id != m_raft_server_id) peer_ids.emplace_back(peer_id);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fetch data might fail and then generate some garbage here if the selected peer doesn't have data in the following case, but anyway, raft will retry:
T1: leader push blob=100 to F1 and F2
T2: F1 and F2 cannot alloc blk because the related shard not committed
T3: F1 append log, fetch data from F2
T4: F2 append log, fetch data from F1
then F1 and F2 will failed to get data, waiting raft retry and select the leader

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks this is a good point.

Maybe we should only fetch data when we are in resync mode ? in that case (only consider 3 copies) leader and another follower both committed.

Lets track this and verify in SH testing, and decide based on metrics,

  • how much improvement we see in fetch latency ?
  • how many more garbages this approach create vs previous?

@JacksonYao287

Copy link
Copy Markdown
Member Author

I find a read verification issue in storage hammer testing. I need to confirm whether it is caused by recent PRs. so I will hold off on merging the PR until I figure out what happens. cc @Besroy @xiaoxichen

@JacksonYao287
JacksonYao287 force-pushed the retry-fetch-data branch 3 times, most recently from 86e32e2 to 5218d1c Compare October 28, 2025 06:21
@JacksonYao287
JacksonYao287 deleted the retry-fetch-data branch July 23, 2026 02:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants