Cut command: pull one column out of a text file
A log file, an export from a billing system, a list of paths from find. The data is already on disk and only one column of it matters. The cut command is the shortest route from that file to that column, and on a Mac it is already installed, which is why it keeps turning up in half-remembered one liners copied out of an old answer.
The trouble starts one step later. The command runs, produces something, and the something is wrong: blank lines where data should be, whole lines passed through untouched, a column that stops halfway through a Japanese filename. None of that is a bug. Each case is a documented behaviour of a tool that is narrower than it looks. Knowing where the edges are is the difference between using cut and fighting it.
Three different things cut can count
The manual page describes one utility with three mutually exclusive modes, and picking the wrong one accounts for a large share of surprising output.
With -b, the list refers to byte positions. With -c, it refers to character positions. With -f, it refers to fields separated by a delimiter. Column and field numbering starts from 1, not from 0, which catches anyone arriving from a programming language.
The list itself is more flexible than most examples show. It is a comma or whitespace separated set of numbers and ranges. A range is two numbers with a dash between them, and both ends are inclusive. A number preceded by a dash means everything from the first field up to that number. A number followed by a dash means everything from there to the end of the line. So -f 3 is one field, -f 2,5 is two, -f 2-4 is three, -f -3 is the first three, and -f 4- is everything from the fourth onward.
Two behaviours in that same paragraph of the manual are worth committing to memory, because they quietly defeat attempts to use cut for reordering. Numbers and ranges may be repeated, overlapping, and given in any order, but if a field is specified more than once it still appears only once in the output. And the output follows the order of the input, not the order written in the list. cut -d , -f 3,1 does not swap two columns; it prints column 1 then column 3. Reordering columns is not something cut does, and no combination of options makes it do so.
The last of those behaviours is the friendly one. Selecting a column that does not exist on a given line is explicitly not an error. A file with ragged rows will not stop the command halfway through; short lines simply contribute nothing for the missing fields.
The delimiter is where most attempts fail
By default -f splits on a tab character. That default is correct for output from many command line tools and wrong for almost every file a person has edited by hand.
-d replaces the delimiter, and the important limitation is in the name: delim is a single character. There is no option for a multi character separator, and no option to treat the delimiter as a pattern. A file separated by a comma followed by a space cannot be cut on ", ". A file where columns are lined up with runs of spaces cannot be cut on "however many spaces there happen to be", at least not with -d.
Then there is the behaviour that produces the most confused questions. When a line contains no delimiter at all, cut passes it through unmodified by default. A header comment, a blank line, a stray note at the bottom of an export: all of them appear in the output verbatim, sitting next to the extracted column and looking like corrupted data. The -s option suppresses those lines instead, and on any real file it is closer to what was wanted than the default. It is worth treating -s as part of the normal invocation rather than as an advanced flag.
The macOS version adds one option that helps with the aligned columns problem. -w uses whitespace, meaning spaces and tabs, as the delimiter, and consecutive spaces and tabs count as a single field separator. That single sentence solves the case that -d ' ' mangles, where two adjacent spaces would otherwise create an empty field between them. The manual is explicit that -w is an extension to the specification, and -w cannot be combined with -d; the synopsis shows them as alternatives.
What no version of cut handles is quoting. A comma inside a quoted field in a CSV file is a delimiter as far as cut is concerned, and there is no flag that changes that. Any file where a text column might contain the delimiter needs a tool that understands the format, not a tool that counts separators.
The version on a Mac is not the version in most examples
Search results for this command are dominated by pages describing the coreutils build that ships with Linux distributions. The manual page on macOS describes a different implementation, and its synopsis is the contract.
That synopsis lists exactly three forms and a small set of options: -b with an optional -n, -c, and -f with -w or -d plus an optional -s. Nothing else. The standards section states that the utility conforms to IEEE Std 1003.2, the POSIX.2 standard, and that -w is the local addition on top of it. An option seen in an example and absent from that synopsis will be rejected, usually with a usage message rather than an explanation.
This is worth a habit rather than a memorised list, because the list changes between systems and between releases. Before adapting a one liner found online, man cut on the machine that will run it settles the question in seconds, and the OPTIONS section is short enough to read in full. The same check applies to every small utility in a pipeline: the behaviour described in a tutorial is the behaviour of whichever build the author happened to have.
The exit status is documented and useful in scripts. The utility exits 0 on success and greater than 0 if an error occurs. Because a missing field is not an error, a non zero status from cut points at the invocation or the file, not at the shape of the data.
Bytes, characters, and text that is not English
The distinction between -b and -c exists entirely for multibyte text, and on a Mac with Japanese filenames or any non ASCII content it decides whether the output is readable.
-b counts bytes. In UTF-8 a Japanese character occupies three bytes, so cutting a fixed number of bytes out of a line of Japanese will usually land in the middle of a character and produce broken output. -c counts characters, which is almost always the intent when a human says "the first ten characters".
When byte positions genuinely are the requirement, for example against a fixed width record format, -n exists to stop the damage. The manual describes it precisely: do not split multibyte characters, and characters are only output if at least one byte is selected and, after a prefix of zero or more unselected bytes, the rest of the bytes forming that character are selected. In practice -n turns a byte range into "whole characters that fall inside this byte range".
Locale settings are part of this. The environment section of the manual names LANG, LC_ALL and LC_CTYPE as affecting execution. A script that runs correctly in an interactive shell and incorrectly from a scheduled job is often a locale difference rather than a logic difference, because the character type rules that -c and -n depend on come from the environment.
When the answer is not cut
Several small utilities overlap with cut, and each has one thing it does that cut cannot. The comparison below is drawn from the manual pages of the versions that ship with macOS.
| Tool | Splits on | Can reorder columns | Reads standard input | Notable limit |
|---|---|---|---|---|
cut |
one character, or whitespace runs with -w |
No | Yes | no multi character delimiter, no quoting |
awk |
a regular expression, set with -F |
Yes | Yes | manual notes only eight bit character sets are handled correctly |
colrm |
character positions only | No | Yes, only standard input | tabs advance the count to the next multiple of 8 |
paste |
joins files rather than splitting lines | Not applicable | Yes, with - |
output separator characters are reused circularly |
tr |
individual characters | No | Yes | translates, deletes or squeezes, never selects fields |
The row that matters most is awk. Its -F option defines the input field separator as a regular expression, which removes the single character limit outright, and fields are addressed as $1, $2 and so on, with $0 for the whole line, so writing them in a different order is trivial. Where cut refuses to reorder, awk reorders by construction. The cost is a second language to hold in your head, and the manual's own BUGS section warns that only eight bit character sets are handled correctly, which matters for the same Japanese text that -c was introduced to protect.
colrm deserves a mention because it is the tool people actually want when they reach for cut -c against terminal output. It removes columns rather than selecting them, it reads only standard input, and its manual is explicit that a tab increments the column count to the next multiple of eight while a backspace decrements it by one. That tab rule is invisible until it is not.
paste is the inverse operation and pairs naturally with cut: extract two columns separately, then join them with a chosen separator. Its -d list is used circularly, so a single option can alternate separators across several joins.
Two examples worth copying exactly
The manual page ends with two examples, and both are better starting points than most of what a search returns because they are guaranteed to work on the machine in front of you.
The first extracts login names and shells from the system password file as name:shell pairs, using a colon as the delimiter and selecting fields 1 and 7. It is the canonical demonstration of -d and a non contiguous field list.
The second pipes who into cut and selects two character ranges, 1-16 and 26-38, to show the names and login times of the users currently logged in. It is also a quiet lesson about -c: the ranges work because that output is fixed width, and they would break the moment the format changed. Fixed width extraction is fragile by nature, which is a reason to prefer -f with a real delimiter whenever the data has one.
Both examples read a file or a pipe and write to standard output, which is the whole shape of the tool. If no file is given, or the file argument is a single dash, cut reads from standard input, so it slots into a pipeline without ceremony.
What to change first
Read man cut on the machine that will run the command, then add -s to any -f invocation and decide between -c and -b before writing the list. Those three steps remove most of the wrong output people attribute to the command itself. If the work involves finding the file, running the command, and checking the result across separate windows, a file manager with a built in terminal removes the switching rather than the typing, and Atriens is built around that single window; the FAQ covers what it does and does not replace.
Frequently asked questions
Why is cut printing whole lines instead of just the column I asked for?
Those lines contain no instance of the delimiter, and the default behaviour is to pass such lines through unmodified. Add -s to suppress lines with no field delimiter characters. Header comments and blank lines are the usual culprits.
How do I use a comma followed by a space as the delimiter?
You cannot. The -d option takes a single delimiter character, and there is no multi character or pattern based equivalent. Use awk -F', ' instead, since -F accepts a regular expression as the field separator.
Can cut swap two columns around?
No. Fields are written in the order they appear in the input regardless of the order in the list, and a field named twice still appears only once. Reordering requires awk, where fields are referenced individually as $1, $2 and so on.
Why does cut break Japanese text in half?
Almost certainly because -b is being used where -c was meant. -b counts bytes and a Japanese character takes several of them, so a byte boundary can fall inside a character. Use -c for character positions, or keep -b and add -n, which prevents multibyte characters from being split.
An option from a tutorial gives a usage error. Is the command broken?
No, the option does not exist in the version installed. The macOS manual page lists the complete set in its synopsis and states that the utility conforms to POSIX.2 with -w as a local extension. Anything outside that synopsis belongs to a different build.