We have a huge BPMN process model with multiple activities. One of the activities was continuously polling an external system every 5 minutes without any time limit. Later, we found that this polling was expensive and the external system had limits, so we changed the model to a time-bound approach.
Now, we need to remove the old process instances using the API. We suspended all the old process instances, so the external system calls from those instances have stopped. However, when we tried deleting the process instances one by one using the REST API, it took a long time. Sometimes deleting a single process instance takes 8 minutes, 20 minutes, or even longer. The process instance deletion is removing both runtime and history data synchronously.
Why does deleting a single process instance take so much time? How can we improve the deletion performance?
Note: The BPMN model has around 20 activities, and each process instance stores around 60 variables, including binary variables. We initially tried deleting the running process instances directly, but we got optimistic locking exceptions. We then tried suspending the process instances first, which worked and stopped the external polling. However, deleting the suspended instances is still taking a very long time.
I am a bit confused, we do not have a public Java API that would delete the runtime and historic instances in one call. To achieve this you’ll need to first delete the runtime and then delete the historic.
Can you provide some more information about how large your process instance is. Having 20 activities is no too much. How many activity instances do you have per process instance. You mentioned that you had a polling every 5 minutes without any time limit, so I would expect that you have a large amount of child activity instances.
Can you also share which Flowable version are you using? Were you using Multi Instance etc.
My guess is amount of data in the history and runtime.
Even if history data is low, if the process instance is polling external system each 5’’ in the loop the execution can still generate a lot of data in ACT_RU_ACTINST table. Check the size of this table first. Delete rows related to your process definition (Should be easy because there is process definition column). The delete api for runtime should work fine after that.
Got it we handled both runtime and history delete API’s in one rest call. And got the historic activity instances “act_hi_actinst” count around 200k huge for the instance.
Is it the sole reason for delete taking time, How can we handle this case then/ How to improve the deletion process.
Note: we are using flowable 7.0.1
I would advise you that you do the runtime deletion and historic deletion separately and try using the HistoricProcessInstance#deleteSequentiallyUsingBatch for more optimized deletion.
This is a release from January 2024, and we have done different improvements in the deletion logic. You might want to upgrade as well.
From flowable 7.0.1 is there any API available to delete process instance asynchronously.
Is it a good approach to delete runtime process instances and will initiate delate historic process instance as async?
Hi @martin.grofcik , Thank you,
I’ve checking ACT_RU_ACTINST table got only few records for the process instance.
If we do two transactions instead of one like delete runtime instances first if done do history delete, But entire delete call as a whole will take the same time right or will it improve?
Hi @filiphr ,
Thanks for your help.
Does this API support Process Instance ID parameter also so that will use process instance Id similar to Process Definition ID. So will consume this API in single process deletion case.
Or Do we have alternative APIs to use for signle deletion.
Is deleteSequentiallyUsingBatch will handle the failed scenarios, i.e if while deleting process instance history got any issue will this retries (3 time by default), will it moves to deadletter job after retries completed.
What is the behaviour currently if we get any issue in async delete of process instance.