Empresas
Empregos
  • Sobre nós
  • Soluções
    • Publicação de vagas
      Publique sua vaga e receba candidatos qualificados em 48h.
    • Avaliações de candidatos
      Mais de 500 testes técnicos e psicológicos, mais anti-fraude.
    • Headhunting
      Busca executiva personalizada do início ao fim.
    • Folha de Pagamento + EOR
      Dispersão de folha e EOR em mais de 15 países da LATAM.
  • Preços
  • Empregos

0

208
Visualizações
How can I display unique words contained in a Bash string?

I have a string that has duplicate words. I would like to display only the unique words. The string is:

variable="alpha bravo charlie alpha delta echo charlie"

I know several tools that can do this together. This is what I figured out:

echo $variable | tr " " "\n" | sort -u | tr "\n" " "

What is a more effective way to do this?

over 4 years ago · Santiago Trujillo
3 Respostas
Responde à pergunta

0

Use a Bash Substitution Expansion

The following shell parameter expansion will substitute spaces with newlines, and then pass the results into the sort utility to return only the unique words.

$ echo -e "${variable// /\\n}" | sort -u
alpha
bravo
charlie
delta
echo

This has the side-effect of sorting your words, as the sort and uniq utilities both require input to be sorted in order to detect duplicates. If that's not what you want, I also posted a Ruby solution that preserves the original word order.

Rejoining Words

If, as one commenter pointed out, you're trying to reassemble your unique words back into a single line, you can use command substitution to do this. For example:

$ echo $(echo -e "${variable// /\\n}" | sort -u)
alpha bravo charlie delta echo

The lack of quotes around the command substitution are intentional. If you quote it, the newlines will be preserved because Bash won't do word-splitting. Unquoted, the shell will return the results as a single line, however unintuitive that may seem.

over 4 years ago · Santiago Trujillo Relatório

0

You may use xargs:

echo "$variable" | xargs -n 1 | sort -u | xargs
over 4 years ago · Santiago Trujillo Relatório

0

Note: This solution assumes that all unique words should be output in the order they're encountered in the input. By contrast, the OP's own solution attempt outputs a sorted list of unique words.

A simple Awk-only solution (POSIX-compliant) that is efficient by avoiding a pipeline (which invariably involves subshells).

awk -v RS=' ' '{ if (!seen[$1]++) { printf "%s%s",sep,$1; sep=" " } }' <<<"$variable"

# The above prints without a trailing \n, as in the OP's own solution.
# To add a trailing newline, append  `END { print }` to the end 
# of the Awk script.
  • Note how $variable is double-quoted to prevent it from accidental shell expansions, notably pathname expansion (globbing), and how it is provided to Awk via a here-string (<<<).

  • -v RS=' ' tells Awk to split the input into records by a single space.

    • Note that the last word will have the input line's trailing newline included, which is why we don't use $0 - the entire record - but $1, the record's first field, which has the newline stripped due to Awk's default field-splitting behavior.
  • seen[$1]++ is a common Awk idiom that either creates an entry for $1, the input word, in associative array seen, if it doesn't exist yet, or increments its occurrence count.

  • !seen[$0]++ therefore only returns true for the first occurrence of a given word (where seen[$0] is implicitly zero/the empty string; the ++ is a post-increment, and therefore doesn't take effect until after the condition is evaluated)

  • {printf "%s%s",sep,$1; sep=" "} prints the word at hand $1, preceded by separator sep, which is implicitly the empty string for the first word, but a single space for subsequent words, due to setting sep to " " immediately after.


Here's a more flexible variant that handles any run of whitespace between input words; it works with GNU Awk and Mawk[1]:

awk -v RS='[[:space:]]+' '{if (!seen[$0]++){printf "%s%s",sep,$0; sep=" "}}' <<<"$variable"
  • -v RS='[[:space:]]s+' tells Awk to split the input into records by any mix of spaces, tabs, and newlines.

[1] Unfortunately, BSD/OSX Awk (in strict compliance with the POSIX spec), doesn't support using regular expressions or even multi-character literals as RS, the input record separator.

over 4 years ago · Santiago Trujillo Relatório
Responde à pergunta
Encontrar trabalhos remotos

Descubra a nova forma de encontrar um emprego!

melhores empregos
Principais categorias de trabalho
Empresas
Postar vaga Preços Comercial
Jurídico
Termos e Condições Política de privacidade
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Recomende algumas ofertas para mim
Preciso de ajuda