0
votes

I am having a lot of trouble for understanding how, the MergeRecord processor works.

I have a flow that is generating, in 1 hour, about 70 000 flowfiles. I need all these flowfiles to be merged into one unique flowfile. The flowfiles are CSV-like, thus it's kind of easy. What is not easy it doing it in one time.

For now I have 2 MergeRecord processors one after the other.

The first one is configured as the following:

  • Min/Max record : 1000/2000
  • Max bin age : 60s
  • Bin number : 2
  • Run schedule : 10 s

The last is configured as the following:

  • Min/Max record : 70/2000
  • Max bin age : 6000s
  • Bin number : 2
  • Run schedule : 60 s

Each incoming queue is on "single node load balance" strategy.

What I expect:

The first MergeRecord, will merge the record according to :

  1. Min record is reached
  2. If a bin is full
  3. If the bin age is reached

Thus leading me with about 70 merged flow files.

The second MergeRecord, will merge the record according to :

  1. Min record is reached
  2. If a bin is full
  3. If the bin age is reached

Thus leading me to 1 flow file.

What I have:

I tried playing with the processor parameters, but it seems that this processor is "only" lead by the run schedule parameters. I have to set it with a high value if I only what single flowfile at the end of my flow. So that the processor will be triggered only when the queue is full. The issue is that 70 000 flowfiles in a queue often lead me to OOM error or Java Heap size error.

Any tips, on the configuration of the mergeRecord processor?

Thank you,

PS: I am running on a 3 nodes cluster.

1

1 Answers

0
votes

The issue is that 70 000 flowfiles in a queue often lead me to OOM error or Java Heap size error.

MergeRecord reads all of the records into memory and sorts them into output bins. That's why you're seeing this. If your records aren't that big, increasing the amount of memory the JVM gets should be enough.

PS: I am running on a 3 nodes cluster.

I'm not sure how you'd merge all of the results from 3 different nodes into a single flowfile. You might be able to get it done by having one MergeRecord that runs on the primary node only and having the input connection set "single node" for the load balancing settings. Just guessing on that one.

Path of least resistance would be just to have each of the 3 nodes do their merges and push to the output data source and set up downstream processes to not have to worry about if it's 1 or 3 files.