[CCR] Make shard follow tasks more resilient for restarts #37239

martijnvg · 2019-01-08T21:44:39Z

If a running shard follow task needs to be restarted and
the remote connection seeds have changed then
a shard follow task currently fails with a fatal error.

The change creates the remote client lazily and adjusts
the errors a shard follow task should retry.

This issue was found in test failures in the recently added
ccr rolling upgrade tests. The reason why this issue occurs
more frequently in the rolling upgrade test is because ccr
is setup in local mode (so remote connection seed will become stale) and
all nodes are restarted, which forces the shard follow tasks to get
restarted at some point during the test. Note that these tests
cannot be enabled yet, because this change will need to be backported
to 6.x first. (otherwise the issue still occurs on non upgraded nodes)

I also changed the RestartIndexFollowingIT to setup remote cluster
via persistent settings and to also restart the leader cluster. This
way what happens during the ccr rolling upgrade qa tests, also happens
in this test.

Relates to #37231

If a running shard follow task needs to be restarted and the remote connection seeds have changed then a shard follow task currently fails with a fatal error. The change creates the remote client lazily and adjusts the errors a shard follow task should retry. This issue was found in test failures in the recently added ccr rolling upgrade tests. The reason why this issue occurs more frequently in the rolling upgrade test is because ccr is setup in local mode (so remote connection seed will become stale) and all nodes are restarted, which forces the shard follow tasks to get restarted at some point during the test. Note that these tests cannot be enabled yet, because this change will need to be backported to 6.x first. (otherwise the issue still occurs on non upgraded nodes) I also changed the RestartIndexFollowingIT to setup remote cluster via persistent settings and to also restart the leader cluster. This way what happens during the ccr rolling upgrade qa tests, also happens in this test. Relates to elastic#37231

martijnvg · 2019-01-09T06:22:37Z

run the gradle build tests 2

dnhatn · 2019-01-09T06:50:40Z

@martijnvg Nice find :) and LGTM but I think it's a good time to define a dedicated exception for the remote client errors.

martijnvg · 2019-01-09T07:27:40Z

but I think it's a good time to define a dedicated exception for the remote client errors.

Agreed. Should we do that in a followup change or as part of this PR?

dnhatn · 2019-01-09T09:14:37Z

Agreed. Should we do that in a followup change or as part of this PR?

A follow-up should be fine :).

…ter_connection_issues

dnhatn

LGTM.

…ter_connection_issues

If a running shard follow task needs to be restarted and the remote connection seeds have changed then a shard follow task currently fails with a fatal error. The change creates the remote client lazily and adjusts the errors a shard follow task should retry. This issue was found in test failures in the recently added ccr rolling upgrade tests. The reason why this issue occurs more frequently in the rolling upgrade test is because ccr is setup in local mode (so remote connection seed will become stale) and all nodes are restarted, which forces the shard follow tasks to get restarted at some point during the test. Note that these tests cannot be enabled yet, because this change will need to be backported to 6.x first. (otherwise the issue still occurs on non upgraded nodes) I also changed the RestartIndexFollowingIT to setup remote cluster via persistent settings and to also restart the leader cluster. This way what happens during the ccr rolling upgrade qa tests, also happens in this test. Relates to #37231

Relates to #37231

martijnvg added >bug v7.0.0 :Distributed Indexing/CCR Issues around the Cross Cluster State Replication features v6.7.0 labels Jan 8, 2019

martijnvg requested a review from dnhatn January 8, 2019 21:44

Merge remote-tracking branch 'es/master' into ccr_improve_remote_clus…

64aaf48

…ter_connection_issues

dnhatn approved these changes Jan 9, 2019

View reviewed changes

martijnvg added 2 commits January 10, 2019 08:43

Merge remote-tracking branch 'es/master' into ccr_improve_remote_clus…

cf49d77

…ter_connection_issues

Merge remote-tracking branch 'es/master' into ccr_improve_remote_clus…

a67fa32

…ter_connection_issues

martijnvg merged commit df48872 into elastic:master Jan 10, 2019

martijnvg added v6.6.0 and removed v6.7.0 labels Jan 10, 2019

This was referenced Jan 10, 2019

[CCR] CCRIT.testAutoFollowing is failing #37231

Closed

LocalIndexFollowingIT#testRemoveRemoteConnection fails on master #37014

Closed

martijnvg added a commit that referenced this pull request Jan 11, 2019

Unmuted test now that #37239 has been merged and backported.

37493c2

Relates to #37231

colings86 added v7.0.0-beta1 and removed v7.0.0 labels Feb 7, 2019

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[CCR] Make shard follow tasks more resilient for restarts #37239

[CCR] Make shard follow tasks more resilient for restarts #37239

martijnvg commented Jan 8, 2019

martijnvg commented Jan 9, 2019

dnhatn commented Jan 9, 2019

martijnvg commented Jan 9, 2019

dnhatn commented Jan 9, 2019 •

edited

Loading

dnhatn left a comment

[CCR] Make shard follow tasks more resilient for restarts #37239

[CCR] Make shard follow tasks more resilient for restarts #37239

Conversation

martijnvg commented Jan 8, 2019

martijnvg commented Jan 9, 2019

dnhatn commented Jan 9, 2019

martijnvg commented Jan 9, 2019

dnhatn commented Jan 9, 2019 • edited Loading

dnhatn left a comment

Choose a reason for hiding this comment

dnhatn commented Jan 9, 2019 •

edited

Loading