Every line below was run, not typed
Run here on 2026-10-03 with GNU Awk 5.3.2 (/usr/bin/awk on Kali) against three small files you can recreate from the next section. Where mawk or BSD awk behave differently, the line says so. For grep patterns, the Grep Pattern Builder builds and tests them in the browser; for whole-pipeline text work, the bash text processing guide puts awk next to grep, sed and sort.
awk cheat sheets are usually a PDF or a table of syntax. This one is the command, the output it actually printed, and one line on why. Copy any line, swap the file name, and it does what the output shows.
What Files Do These Examples Use?
A web server log with documentation-range IP addresses, a /etc/passwd-format file, and a CSV with a header:
In access.log, field 1 is the client IP, $7 the path, $9 the status code and $10 (also $NF) the response size in bytes. Three two-line scratch files appear later: a.txt holds alpha 3, beta 7, gamma 2; b.txt holds delta 9, epsilon 1; c.txt is a copy of a.txt.
How Do I Print Columns?
Fields split on runs of whitespace; $1 is the first.
The comma inserts the output separator (a space). Without it, $1 $9 would glue the fields together.
NF is the field count, so $NF is the last field and $(NF-1) the one before, whatever the line length. The parentheses matter: $NF-1 is the last field minus one.
-F sets the input separator. NR > 1 skips the CSV header.
OFS is the output separator, set in BEGIN before the first line is read. Tab-separated output pastes cleanly into a spreadsheet.
printf for aligned columns and fixed decimals. Unlike print, it adds no newline: the \n is yours to write.
How Do I Filter Rows?
A pattern with no action prints the whole matching line. The comparison is numeric because both sides look like numbers.
/re/ matches anywhere in the line, like grep. $7 ~ /re/ matches one field only, which is what stops /api in a query string or user agent from matching.
Human accounts (UID 1000 and up on most distros), and accounts that can log in. !~ is "does not match".
Line numbers and ranges by NR.
-v passes a shell value in as an awk variable. Splicing "$LIMIT" into the program text instead breaks on spaces and quotes, and lets the value inject awk code.
How Do I Sum, Count and Find the Maximum?
Variables start at 0; END runs after the last line, when NR is the total line count.
Arithmetic on fields per row, summed.
Arrays keyed by a field are awk's GROUP BY. for (k in arr) visits keys in no fixed order, hence the sort. The second line is the top-talkers report for any log.
Deduplicate while keeping first-seen order, which sort -u cannot do. seen[$1]++ is 0 (false) the first time a key appears, so ! makes it true exactly once.
The row with the largest value, and a line count (wc -l without the file name).
What Do NF, NR, FNR and FILENAME Tell Me?
The same file has 1 field per line with the default separator (no spaces in root:x:0:…) and 7 with -F:. When a column comes out empty, print NF first: the separator is usually the problem.
FNR restarts per file; NR would have run 1 to 5.
| Variable | Meaning |
|---|---|
$0 | the whole line |
$1…$NF | the fields |
NF | number of fields on this line |
NR | line number across all input |
FNR | line number within the current file |
FILENAME | current input file |
FS / -F | input field separator (default: runs of whitespace) |
OFS | output field separator (default: one space) |
BEGIN {…} | runs before the first line |
END {…} | runs after the last line |
How Do I Change Text With awk?
toupper, split into an array, and sub (first match; gsub for all). After sub changes $0, awk re-splits the fields.
Assigning to a field rebuilds $0 with OFS between fields, which also collapses any original spacing. That surprises people when they expected the rest of the line untouched.
1 is a pattern that is always true, with the default action (print). In-place editing is GNU awk only; elsewhere write to a temporary file and mv it. Never awk … file > file: the shell empties the file before awk reads it.
When Is awk the Wrong Tool?
cut splits on every single space, so two spaces make an empty field 2. awk treats runs of whitespace as one separator, which is why it is right for ps, df and ls -l output and cut is not. The other direction:
Plain -F, splits inside quotes. gawk 5.3 added --csv, which parses quoted fields properly; on older awks, use a real CSV tool (mlr, Python's csv). And for a pure find-and-replace on lines, sed is shorter; for "does this line contain X", grep is faster. awk earns its place when you need fields, arithmetic or a total.
Prerequisites
Any awk (gawk, mawk, BSD awk, busybox) runs everything above except -i inplace and --csv, which need GNU awk (5.3+ for --csv). Check with awk --version; mawk answers mawk -W version.
Frequently Asked Questions
How do I print a specific column with awk?
Use awk '{print $N}' file, where N is the column number: awk '{print $1}' prints the first, awk '{print $1, $3}' prints the first and third separated by a space, and awk '{print $NF}' prints the last whatever the line length. Columns split on runs of spaces and tabs by default. For another separator, set it with -F: awk -F: '{print $1}' /etc/passwd prints user names.
How do I sum a column with awk?
Accumulate in a variable and print it in an END block: awk '{sum += $3} END {print sum}' file. Variables start at zero, so no initialisation is needed. Skip a header with NR > 1, and use printf for decimals: awk 'NR > 1 {s += $3} END {printf "%.2f\n", s}' file. For an average, divide by NR (or by NR - 1 when you skipped a header).
What is the difference between NR and FNR in awk?
NR counts lines across all input files and never resets. FNR counts lines within the current file and restarts at 1 for each new file. With one file they are equal. With several, FNR == 1 marks the first line of each file, and NR == FNR is true only while awk reads the first file, which is the standard idiom for loading one file into an array before processing a second.
Why does cut give the wrong column when awk gives the right one?
cut -d' ' treats every single space as a separator, so two spaces in a row create an empty field and shift every column after it. awk's default separator is a run of spaces and tabs, and leading whitespace is ignored, so aligned output from ps, df or ls -l splits the way it looks. Use cut for files with exactly one delimiter character between fields, such as CSV without quotes, and awk for anything space-aligned.
Can awk edit a file in place?
GNU awk can, with gawk -i inplace 'program' file, which rewrites the file with awk's output. mawk and BSD awk cannot: mawk rejects -i as not an option. Portably, write to a temporary file and move it over the original: awk 'program' file > file.tmp && mv file.tmp file. Never redirect awk's output to the file it is reading, which empties the file before awk reads it.
The one-liners above are the building blocks of most log and report scripts; The Production Bash Toolkit ships ShellCheck-clean scripts that use them, and Read a File Line by Line covers the cases where a bash loop is clearer than awk.
Part of the bash snippets collection
Related Scripts
- Find and Replace With sed — line edits, where sed is shorter than awk
- Search Files for Text With grep — finding the lines before awk splits them
- Read a File Line by Line — the bash loop, and when it beats awk
- Bash String Manipulation — splitting and trimming without spawning awk