Audit Cache/Store (ie Replay and Recovery)
-
bahuvrihi
When an error occurs:
- set termination/raise termination error
- mark the failing task(s)
- as you unwind, cache the results for each audit
- requeue the lead tasks
Then during debuggin'
- the failed task(s) located and manipulated
- when the app is restarted the tasks run with cached results, allowing the join state to be recreated
A similar thing can be accomplished if results are cached along the way, for example to an external data store. How caching/auditing occurs is a separate issue with speed vs memory issues.
-
bahuvrihi
One issue is that this may run into a duplication problem if a join has enqued tasks. If:
a -- b -- c --[0][1,2]qAnd task c fails, I guess this is actually ok because by the time c runs, b will have already been taken off the stack.
Not ok, however if:
a -- b --:q c -- d --[0][1,3]Because c will be enqued after b, then d fails. Replay will enque d c again, leading to duplication. If b doesn't run at all, however, then this model is ok.
So, this would have to be a dispatch-level cache/replay and not a middleware thing. Don't dispatch if completed.
Please Sign in or create a free account to add a new ticket.
With your very own profile, you can contribute to projects, track your activity, watch tickets, receive and update tickets through your email and much more.