Surface per-task status summary in Helix JobContext#214
Open
LZD-PratyushBhatt wants to merge 1 commit into
Open
Surface per-task status summary in Helix JobContext#214LZD-PratyushBhatt wants to merge 1 commit into
LZD-PratyushBhatt wants to merge 1 commit into
Conversation
LZD-PratyushBhatt
requested review from
PranaviAncha,
arkmish,
kabragaurav,
laxman-ch,
ngngwr,
sjainit and
thestreak101
as code owners
July 21, 2026 04:17
When a Task Framework job is configured with a high FailureThreshold so that every partition task is allowed to run to completion, the job's own status flag stays COMPLETED even if individual tasks fail. That masks partial failures: operators watching only the job state cannot tell that some partitions errored, aborted, or timed out. This adds an aggregated, job-level task status summary that the controller computes and writes into the JobContext when a job reaches a terminal state (completed, failed, or timed out). The summary carries completed/failed counts, a per-state breakdown, and the list of failed partitions, so partial failures stay visible even when the job flag is COMPLETED. Reading the summary is a clean alternative to the Integer.MAX_VALUE FailureThreshold workaround. helix-core: - JobContext#updateTaskStatusSummary computes the summary from per-partition states and stores it as the TASK_STATUS_SUMMARY simple field (JSON); JobContext#getTaskStatusSummary reads it back. - AbstractTaskDispatcher invokes it at the three terminal choke points (markJobComplete, failJob, handleJobTimeout). helix-front: - Job detail view gains a "Task Summary" tab that parses TASK_STATUS_SUMMARY and highlights failures. Tests: - TestJobTaskStatusSummary integration test (job COMPLETED, summary reports the failures). - JobTaskSummaryDriver standalone end-to-end driver (runs against a live ZK). - job-detail component specs for the new getters. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
LZD-PratyushBhatt
force-pushed
the
lzd/job-task-status-summary
branch
from
July 21, 2026 04:45
7edef10 to
5d3303d
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When a Task Framework job is configured with a high FailureThreshold so that every partition task is allowed to run to completion, the job's own status flag stays COMPLETED even if individual tasks fail. That masks partial failures: operators watching only the job state cannot tell that some partitions errored, aborted, or timed out.
This adds an aggregated, job-level task status summary that the controller computes and writes into the JobContext when a job reaches a terminal state (completed, failed, or timed out). The summary carries completed/failed counts, a per-state breakdown, and the list of failed partitions, so partial failures stay visible even when the job flag is COMPLETED. Reading the summary is a clean alternative to the Integer.MAX_VALUE FailureThreshold workaround.
helix-core:
helix-front:
Tests:
Locally tested in real helix cluster:

Issues
(#200 - Link your issue number here: You can write "Fixes #XXX". Please use the proper keyword so that the issue gets closed automatically. See https://docs.github.com/en/github/managing-your-work-on-github/linking-a-pull-request-to-an-issue
Any of the following keywords can be used: close, closes, closed, fix, fixes, fixed, resolve, resolves, resolved)
Description
(Write a concise description including what, why, how)
Tests
(List the names of added unit/integration tests)
(If CI test fails due to known issue, please specify the issue and test PR locally. Then copy & paste the result of "mvn test" to here.)
Changes that Break Backward Compatibility (Optional)
(Consider including all behavior changes for public methods or API. Also include these changes in merge description so that other developers are aware of these changes. This allows them to make relevant code changes in feature branches accounting for the new method/API behavior.)
Documentation (Optional)
(Link the GitHub wiki you added)
Commits
Code Quality
(helix-style-intellij.xml if IntelliJ IDE is used)