In the past, my client would call to tell me that a problem occurred in their software. I would log in to the site, and look at the error logs.
However, I found that error logs tend to explain WHAT error occurred but not WHY the error occurred. An understanding of what led to the issue requires previous state information, which is only contained in DEBUG logs.
So almost every time, I would have to change the log level, restart the software, and spend a lot of time trying to recreate the issue.
I decided to leave the production code running in DEBUG log level, but with one adjustment:
I capped the max journal size using journald.conf to 10GB. On a 500GB machine this seemed fine to me.
Now I can use journalctl --since and journalctl --until to filter the huge log to the time period when my client said the error occurred.
And now I don't waste time re-creating the issue when problems come up.
My question:
What are the implications of leaving production code running on a client's site in a verbose DEBUG level?
I found the answer here inadequate: Log levels in Production
