I am cleaning a string variable in Stata that has numeric values but occasionally has values formatted as a range, as in 1-50 or 1-3, etc.
When I try to destring these variables, these pesky ranges prevent me from doing so.
What I would like to do is replace the range with the average of the first number and the last number in the range. I have tried the following string functions to do this:
replace `var' = ((regexs(1) + regexs(3))/2) if regexm(`var', "([0-9]*)([\-])([0-9]*)")
However, Stata cannot understand the average ((regexs(1) + regexs(3))/2) because it reads regexs(1) and regexs(2) as substrings.
I know I could do this by creating new variables, but the data I am working with has thousands of variables, so I would really prefer to just replace the existing string.
Any ideas on how to do this?
Thanks in advance