Awk example lines you can reuse on your own files
Searching for an awk example almost always lands on a page of one-liners lifted from a Linux tutorial. Most of them run on a Mac exactly as printed. A few produce output that looks plausible and is wrong, and the reason is never mentioned on the page the line came from. The awk installed on macOS is not the awk those tutorials were written against, and the difference shows up precisely where the data stops being plain ASCII.
That is the useful thing to sort out before collecting more one-liners. What follows separates the lines that are safe to paste from the ones that need checking, and names the specific place where the shipped version stops being reliable.
The awk on a Mac is not the awk in the examples
Open the manual page and the first line reads awk - pattern-directed scanning and processing language. The page is dated 2020-11-24 and its SEE ALSO section points at a single book: A. V. Aho, B. W. Kernighan, P. J. Weinberger, The AWK Programming Language, Addison-Wesley, 1988. This is the original implementation, the one usually called the one true awk, and it is a much smaller program than the GNU version that most tutorials assume.
The consequence is written down in the manual, in the BUGS section at the very bottom, where almost nobody reads:
Only eight-bit characters sets are handled correctly.
That single sentence decides whether a given example can be trusted. Anything that counts characters, slices a string by position, or changes case is operating on bytes. A Japanese filename, a name with an accent, a currency symbol, a smart quote pasted in from a document: all of them occupy more than one byte, and length, substr, toupper and tolower will count and cut in the wrong places. Nothing errors. The numbers just come out larger than expected and the substrings come out truncated mid-character.
Lines that only split on a separator and print whole fields are unaffected, because whole fields are copied byte for byte. That covers the large majority of practical use, which is why the problem stays hidden for years and then surfaces on one specific file.
The GNU implementation is a separate package. Homebrew carries it as gawk, currently at version 5.4.1, and the formula notes that it installs under the name gawk, not awk. To have awk mean the GNU version, the formula's own instructions are to put its gnubin directory ahead of everything else in PATH:
PATH="$HOMEBREW_PREFIX/opt/gawk/libexec/gnubin:$PATH"
| What to check | awk as shipped with macOS | gawk from Homebrew |
|---|---|---|
| Command name | awk |
gawk, or awk when gnubin is added to PATH |
| Version | part of the operating system | 5.4.1 |
| Text outside ASCII | manual states only eight-bit character sets are handled correctly | not covered by that manual page |
| Other binaries installed | none | awk, gawk, gawk-5.4.1, gawkbug |
| Homebrew installs, last 30 days | not applicable | 1,477 |
Worth noting how small that install figure is. Over the same 30 days the git formula records 37,186 installs. Reaching for the GNU version is not the common path, which is another reason the byte-counting limitation stays undiscovered.
There is one more switch on the shipped version that matters for anyone copying substitution examples. The manual says that when POSIXLY_CORRECT is set in the environment, awk follows the POSIX rules for sub and gsub with respect to consecutive backslashes and ampersands. An ampersand in a replacement string stands for the matched text, and a backslash escapes it, so a replacement containing either one can behave differently depending on whether that variable happens to be set. Tutorials never mention it because on the machine where they were written it was not set.
Everything awk does is one pattern and one action
Every example that looks cryptic is the same two-part shape, and reading it that way makes the rest mechanical:
pattern { action }
A missing action means print the whole line. A missing pattern means the action runs on every line. That is the entire grammar, repeated with newlines or semicolons between statements.
Fields are named $1, $2 and so on, with $0 standing for the whole line. By default the separator is any run of whitespace, so no configuration is needed for output from ls -l or a space-aligned report. Setting FS to something else, either with -F on the command line or by assigning it inside a BEGIN block, switches to commas, tabs or a regular expression. A detail the manual calls out and most examples omit: if FS is set to the empty string, the line is split into one field per character, which is occasionally the shortest route to a character-by-character pass.
Two special patterns hold the rest together. BEGIN runs before the first line is read, which is where separators get set and headers get printed. END runs after the last line, which is where totals get printed. Both may appear more than once and run in the order awk reads them.
The -v option assigns a variable before the program starts, which is how a value from the shell gets in without quoting gymnastics. Any argument shaped like var=value in the file list is also treated as an assignment, executed at the point where that file would have been opened. That second form is easy to trip over: a file genuinely named count=3 will not be read.
Example lines worth keeping
These are the shapes that come up repeatedly. Four of them are lifted straight from the EXAMPLES section of the manual page, which is the shortest reliable reference available offline.
Print lines longer than 72 characters:
awk 'length($0) > 72' report.txt
Print the first two fields in the opposite order:
awk '{ print $2, $1 }' list.txt
Handle input separated by commas, or tabs, or both, then swap the first two fields:
awk 'BEGIN { FS = ",[ \t]*|[ \t]+" } { print $2, $1 }' data.csv
Add up the first column and print the sum and the average:
awk '{ s += $1 } END { print "sum is", s, " average is", s/NR }' numbers.txt
Print everything between two markers, inclusive. A pattern pair separated by a comma matches the whole range:
awk '/start/,/stop/' log.txt
Count how many times each value appears in a column. Array subscripts can be any string, which makes this a three-word frequency counter:
awk '{ n[$3]++ } END { for (k in n) print n[k], k }' access.log
Print the last field of every line, whatever the line length. NF holds the field count, and $NF reads the field it names:
awk '{ print $NF }' paths.txt
Print the filename alongside a match when several files are passed at once:
awk '/ERROR/ { print FILENAME ":" FNR ": " $0 }' *.log
The last one is the strongest argument for learning FNR. NR keeps counting across every file, while FNR restarts at each new file, so FNR is the number that matches what an editor would show.
Two smaller shapes round out the set. Printing a header before the rest of the output takes nothing but a BEGIN block, and because OFS can be changed in the same place, a report can be turned into tab separated output in one go without touching the rest of the program. And when several files are passed but only the first few lines of each are wanted, nextfile skips the remainder of the current file and moves on, which is considerably faster than reading a large log to the end only to discard it.
The habit worth forming with all of these is to run them first against a copy of the data cut down to twenty or thirty lines. Awk gives no warning when a pattern matches nothing, so a line that prints an empty result looks identical whether the filter is too narrow or the field number is off by one. A small input makes the difference obvious in a second, and the same line then runs unchanged over the full file.
The built-in variables that change the answer
Six of them do most of the work, and knowing them turns guesswork into reading.
NR is the record number counted across all input. FNR is the record number within the current file. NF is the number of fields in the current record, which makes $NF the last field and $(NF-1) the one before it. FILENAME is the file currently being read.
FS splits input into fields, and OFS joins them on output, defaulting to a single space. This pair is where a common surprise lives: changing FS alone does not change the output separator, so a comma-separated file processed with -F, prints back out space separated unless OFS is set as well. Touching any field, even by assigning a field to itself, is what forces the line to be rebuilt with the new OFS.
RS and ORS do the same job one level up, for records rather than fields. Setting RS to the empty string makes blank lines separate records, which parses paragraph-formatted data with no extra logic. Setting RS to more than one character makes it a regular expression, and records are separated by whatever matches.
Three more come up in specific jobs. ARGV and ARGC hold the command line and are assignable. ENVIRON exposes environment variables by name. CONVFMT and OFMT both default to %.6g, which is the reason a long decimal sometimes prints shorter than expected: the number is intact, the format is rounding it.
Where awk stops being the right tool
Two limitations are stated plainly in the manual, and both are worth respecting rather than working around.
The first is type conversion. The BUGS section says there are no explicit conversions between numbers and strings. The documented workarounds are to add 0 to force a number and to concatenate "" to force a string. Skipping that step is how "10" ends up sorting before "9", or how a leading-zero account number silently loses its zero. Anywhere values arrive as text and get compared, one of those two coercions belongs in the line.
The second is scope. The manual's own words are that the scope rules for variables in functions are a botch and the syntax is worse. Function parameters are local; every other variable is global. The only way to get a local variable is to declare extra parameters the caller never passes. For a five-line filter this does not matter. For anything with several functions sharing state, the language is fighting the work, and a script in a language with real scoping will be shorter and easier to fix later.
There is also a practical ceiling around reading and writing. getline pulls the next record, optionally from a named file or from the output of a command through a pipe, returning 1 on success, 0 at end of file and -1 on error. system runs a command and returns its exit status. Both work. Both also mean the one-liner has become a program that shells out, and at that point the pipeline is easier to read than the awk line that hides it.
Checking output instead of trusting it
The failure mode with awk is not an error message. It is a plausible number. So the habit that matters is comparing a small slice of the output against the input by eye before running the line over everything.
That comparison is where the shape of the working environment starts to cost time. The file being processed sits in a folder window, the command runs in a terminal somewhere else, and the result has to be read back against the original. Two applications, one job, and a cd typed out every time the folder changes. A file manager with a built-in terminal removes that specific step by giving each folder its own shell, so the command runs where the files already are. The Features page covers how the folder view and the terminal stay pointed at the same place, and Compared with other file managers is honest about what such a tool does not do: it is not an editor and not a notes application.
The other half of checking is not being at the desk. A pass over a large directory takes minutes, and watching it finish is wasted time. Reading the terminal from a phone, which the From iPhone and iPad page describes, turns that wait into something checked once from elsewhere.
What to change first
Read the BUGS section of the local manual page before collecting any more one-liners, because it names the one limitation that makes correct-looking output wrong. Then decide once whether the data will ever contain text outside ASCII, and install the GNU version if it will. If the checking loop between folder and terminal is the slow part rather than the awk itself, Atriens puts both in one window.
Frequently asked questions
Does an awk example written for Linux work on macOS?
Most do, unchanged, because they only split lines on a separator and print fields. Lines that count characters, cut substrings by position or change case are the exception, since the shipped manual page states that only eight-bit character sets are handled correctly. Data that is plain ASCII behaves the same on both.
How do you tell which awk is actually running?
Check whether a gnubin directory has been placed ahead of the system directories in PATH, since that is what the Homebrew formula instructs for making awk mean the GNU version. Without that line the name awk resolves to the copy that came with the operating system, and the GNU one answers to gawk.
Why does a comma separated file print back out with spaces?
FS controls how input is split and OFS controls how output is joined, and they are independent. Passing -F, changes only the input side, so OFS also has to be set, and at least one field has to be assigned for the record to be rebuilt with it.
When is awk the wrong choice?
When the program needs more than a couple of functions sharing state. The manual itself calls the scope rules for variables in functions a botch, and the only way to make a variable local is to add parameters no caller passes. At that size, a script in a language with real scoping is easier to read and to correct.
Is there a reference available without a network connection?
The manual page installed on the machine is the whole language in a few screens, and its EXAMPLES section carries working lines for printing fields in reverse order, summing a column, filtering by line length and printing everything between two markers. It is shorter and more accurate than most tutorial pages.