When working with text, whether authoring or editing, we often need to change something. Name of some feature from A to B. Spelling error. Numerical value. A phrase that does not fit the rest of text. And usually, we just delete the wrong text and type in the correct one. If a spelling error or incorrect spelling variant is repeated across the document, we may use Find/Change to correct all instances at once.
Things get a bit more difficult when we want to correct more than a single word at a time – but we can use wildcards. There are two commonly used wildcard symbols, and as you probably know, “*” stands for “any string” and “?” means “single character”. So searching for “anaestheti*” will return anaesthetic, anaesthetics, anaesthetised, anaesthetist etc., while searching for “anaestheti?” will return only “anaesthetic” (if “Whole Word” option is checked).
Wildcard-based search may be sufficient in a number of situations, where all you need is to find some text and change its attributes, for example character style. It is especially true in FrameMaker since FM wildcards support some features copied from regular expressions. However, if we need to conduct more advanced searches, conditional searches or perform some kind of operations on the text we find, we should reach for full power of regular expressions (“regex”).
In This Article
- Meet regular expressions
- Alphabeth for Beginners
- Hang on
- Grammar Rule[sz]
- Benefits of Laziness
- Bean Counter
- Numbers Crunching
- What If
- Bring it on
- Reference
Meet regular expressions
What are they? Regular expressions are described as “a sequence of characters that define a search pattern, mainly for use in pattern matching with strings, or string matching, i.e. “find and replace”-like operations” [Wikipedia]. Somewhat like in the case of wildcards, there is a “vocabulary” – set of symbols that can be used to find a string of characters matching our expression, only the symbol set is much bigger – and a special kind of “grammar” which enables the creation of long and complex search patterns. Plus once we find something, we can usually modify it.
To give you some examples:
- Let’s say we want to find all occurrences of “FrameMaker” along with its version number in a FrameMaker Wikipedia article to change its character format. Can we do that with wildcards? Yes, but given the possible variations in numbering (2.0, 5.1.2, 2015) we would need to run the search at least three times. Why waste the time, if we can do it with one regular expression? (
FrameMaker \d+(\.\d+)*) - In Polish texts we often encounter single letter conjunctions – i, o, u, z, a, w – and much like widow lines for paragraphs, it is considered poor typography if a line ends with such single letter. The simplest solution is to replace standard space with non-breaking one for them. Again something that can be done with a single regular expression.
- Find:
(\b[iouzaw]\b) (\w); - Change:
$1 $2– with a non-breaking space on the “Change” side). - Do you need to change date notation from 12/24/2015 to 12-24-2015? Or maybe even from American to European format (24.12.215)? Piece of cake (see below).
- Remove all double spaces in one go? Easy:
- Find
\s{2,} - Change:
\s - Insert (or remove) spaces between numerical values and units? No problem (see below).
- Add a missing word to complex product name in inflected language, but only if it’s really missing? Will do (see below).
So, how do we actually accomplish all these things? You need to learn basic vocabulary and grammar of regular expressions. And you can do this either by reading the introduction below, or just skip to the reference table at the end and rifle through examples. All regular expressions in the text are written with \s{2,} font, and quotation marks around them (if present) are not part of the expressions. And remember – while the expressions may look scary, they are, in fact, quite regular. Once you grasp the meaning of symbols and basic rules, you’ll be able to use them at no time.
Alphabet for Beginners
If we want to find any particular symbol/letter/number using regular expression, we just type that symbol/letter/number. So anaesthetic is a regular expression which matches only that particular sequence of lower-case letters (regular expressions are case-sensitive).
If we need to find any character or symbol, we use period (.). Single dot means single symbol of any kind.
To get somewhat more specific matches we can define character class by using square brackets [ ]. For example [abcd] will match “a”, “b”, “c” or “d”. If you are looking for a symbol within continuous range, instead of defining all elements of class, we can simply state first and last: [a-d], or [0-9] to match any digit. The order in ranges is always the same as in ASCII character table, so to match any letter regardless of case we can use class [A-Za-z].
There is also a special type of class used for negation: if the first symbol within square brackets is the circumflex (^), expression will match anything except the class content. For example [^abcd] will match any single symbol except for a, b, c or d. Class can be seen as logical alternative for set with single symbol elements.
If we need to employ alternatives for longer elements, we can use pipe character “|” to separate elements: Monday|Tuesday|Wednesday|Thursday will match any of these four weekday names. Unfortunately there is no simple way to exclude longer strings for search – matching anything but some string is usually bit tricky.
Do we have to define these classes every time? That depends on the class. There are some shortcuts defined in regular expression vocabulary:
\d= digit =[0-9]\D= digit negation =[^0-9]\w= word character =[A-Za-z0-9_]\W= word negation =[^A-Za-z0-9_]\s= white space (includes all spaces, tabs and end of line characters)\S= white space negation\u####= Unicode character number ####, e.g.\u2212for n-dash (–)
Of course this is just the most basic set. Instead of shortcuts you can also use Unicode categories, which are much more versatile with regard to various scripts and non-ASCII characters – see the reference tables at the end of this post.
Backslash by itself is a special escape character. You know that a period means any symbol, but what do we do if we want to match only period? Simply escape it with backslash: \. . But since backslash is special character too, if we need to match backslash, we have to escape it too: \\
Hang on – anchors
There are two special symbols used to define start (^) and end ($) of the line, plus shortcut defining word boundary (\b) they are called “anchors”. To recall our earlier example, anaesthetic will match that string anywhere in text, but ^anaesthetic will match only, if the string will appear at the beginning of line (paragraph). By itself aesthetic will also match “anaesthetic” and “anaesthetics”, aesthetic\b will match “anaesthetic” but not “anaesthetics” and \baesthetic\b will match neither. It’s a bit like searching with “Whole Word” option, but with more precise control.
Grammar rule[sz]
This was essential “vocabulary”, now we’ll see some “grammar”. Let’s start by introducing operators (quantifiers), since quite often we are interested in some longer strings, not just particular symbols. Operators works for the object immediately to their left. There are three basic operators:
*= 0 or as much as possible. “.*” matches any string. That is, it can match whole paragraph, but also a completely empty paragraph (0 occurrences). “\d*” will match any string of numbers, but also “empty” string, e.g. boundary between letters. You need to be careful with this quantifier.?= 0 or 1..?matches empty string or single symbol of any kind. “\d?” matches no digits or exactly one digit. The concept of “empty” match is a bit tricky, but I suggest to simply enter these expressions into the Find/Change dialog with Regular Expressions option checked and click Find repeatedly to see what will be matched.+= 1 or as much as possible. That one is much more intuitive: “.+” will match any string with at least one symbol and\d+will match any string of digits, but not less than one.
We can use these quantifiers for example to find both spelling versions of an[a]esthetic with a single expression. ana?esthetic will match both anesthetic and an a esthetic, because letter “a” can occur 0 or 1 time.
Benefits of Laziness
When using operators, we need to remember, that by default they behavior is “greedy”, that is, they want to match the longest possible string within paragraph (in FM with default settings) or even whole text.
We can expand our search to cover both spelling variants and different word suffixes. ana?esthe.*\b will match anaesthe or anesthe followed by any letter occurring any number of times until the end of word. Unfortunately, the greedy behavior will make the expression match up to the last word boundary it can find (since period matches also a space):
Now, being eager like that can be a good thing, but as they say, laziness is one of the greatest programmer virtues. So in order to avoid such excessive matching we have to use “lazy” matching operator combinations:
*?= zero or as little as possible+?= one or as little as possible
If we modify our expression accordingly: ana?esthe.*?\b
We can match exactly what we want. The alternative way to safely match only single word in this case would be to use ana?esthe\w+ that way matching will end at the first non-word character (e.g. space or punctuation mark).
Bean Counter
Are these the only quantifiers you have at your disposal? No, for precise matching we can define either the exact number of occurrences we want to match or a range of occurrences using curly brackets {}:
\d{4}= matches exactly 4 digits\d{2,4}= matches from 2 to 4 digits (2, 3, 4)\d{,4}= matches up to 4 digits\d{4,}= matches at least 4 digits
So for our (an)aesthetic example ana?esthe\w{3} will match anaesthesia and anesthetic, but not anesthetist (\w{3} will match exactly three of any “word” characters).
Now it’s time to familiarize ourselves with the use of parentheses ( ). I’ve already written that regular expression allow us to manipulate results of our searches. This is possible, because the match is stored in memory, numbered and can be recalled. Whole match is always numbered “0”. What’s important, we can define parts of search expression to be separate groups, and we do that by putting them inside parentheses. Content of the groups is recalled by $num, e.g. $1 for group number 1.
Numbers Crunching
Let’s say we want to find a date in American format: 12/24/2015. A simple expression for matching this can look like this:
\d\d/\d\d/\d\d\d\d – or, a more elegant and readable version – \d{2}/\d{2}/\d{4}
Match two digits, slash, again two digits, slash and four digits. Now, if we want to somehow manipulate this date, we have to define groups: