freebsd-dev

Author	SHA1	Message	Date
Tim J. Robbins	e5996857ad	Make regular expression matching aware of multibyte characters. The general idea is that we perform multibyte->wide character conversion while parsing and compiling, then convert byte sequences to wide characters when they're needed for comparison and stepping through the string during execution. As with tr(1), the main complication is to efficiently represent sets of characters in bracket expressions. The old bitmap representation is replaced by a bitmap for the first 256 characters combined with a vector of individual wide characters, a vector of character ranges (for [A-Z] etc.), and a vector of character classes (for [[:alpha:]] etc.). One other point of interest is that although the Boyer-Moore algorithm had to be disabled in the general multibyte case, it is still enabled for UTF-8 because of its self-synchronizing nature. This greatly speeds up matching by reducing the number of multibyte conversions that need to be done.	2004-07-12 07:35:59 +00:00
David E. O'Brien	333fc21e3c	Fix the style of the SCM ID's. I believe have made all of libc .c's as consistent as possible.	2002-03-22 21:53:29 +00:00
David E. O'Brien	8fb3f3f682	Remove 'register' keyword.	2002-03-21 18:49:23 +00:00
Daniel C. Sobral	6d902efe43	Since we have modified charjump to be CHAR_MIN-based, we have to correct the offset when we free it. Caught by: phkmalloc	2000-07-08 09:45:17 +00:00
Daniel C. Sobral	a5378e623a	Fix memory leak introduced with regcomp.c rev 1.14.	2000-07-02 15:58:54 +00:00
Rodney W. Grimes	58f0484fa2	BSD 4.4 Lite Lib Sources	1994-05-27 05:00:24 +00:00

6 Commits