In my embedded application, which is very memory sensitive, I noticed some of the newlib functions using a lot of stack space. By looking at the source code of newlib, specifically memmem.c in this case, I noticed two defines, PREFER_SIZE_OVER_SPEED and __OPTIMIZE_SIZE__, which can reduce the memory usage drastically. As far as I understand, these should be defined when compiling newlib to make use of the "optimized for size" libraries. Since I am using a cortex-M3 micro controller, is there any ARM toolchains out there which uses a "optimized for size" newlib or provide the option for using it, or should I try to build it myself. Furthermore, when building newlib, should I also build GCC or can I just build the library and use it with my current toolchain. Currently I am using CoIDE with their supplied toolchain.
2 Answers
You only need build the library, not the compiler.
However I would expect any size optimisation to relate to code size rather than stack size. Stack size would only be reduced if the size or number of auto variables were reduced and generally that is determined by the required functionality not the optimisation of the algorithm.
While it is true that often high-level operations involving the movement of large amounts of data can be speeded by utilising more memory, I would say that such opportunities are minimal at the level of the C standard library, so "prefer size over speed" is all about code size not data memory usage.
You're using memmem which is not a standard function. It's a GNU extension in glibc. The code you're actually running is in str-two-way.h. I didn't study it, but it says it's a sub-linear string search like Boyer-Moore and points you to the wikipedia article on Boyer-Moore. Of course that's going to have some memory costs.
Since it's not even a standard function, there is no reason to use newlib's implementation if you don't like it. Just use your own substring search function. If you know quadratic time is good enough, just use the 5 line loop from memmem.c in you own code. You might want to check that memcmp does the right thing wrt unaligned loads (if your architecture supports those). If it doesn't, a manual nested loop might be faster than calling memcmp.