Empresas
Empleos
  • Sobre nosotros
  • Soluciones
    • Publicación de vacantes
      Publica tu vacante y recibe candidatos calificados en 48h.
    • Evaluación de candidatos
      500+ pruebas técnicas y psicológicas, más anti-fraude.
    • Headhunting
      Búsqueda ejecutiva a la medida de principio a fin.
    • Nómina + EOR
      Dispersión de nómina y EOR en más de 15 países de LATAM.
  • Precios
  • Empleos

0

338
Vistas
Ruby gsub for exact number of words

I have a piece of code, where I can switch words from @post.swap_content to hyperlinks by keyword. For example, I have a word 'michigan' in @post.swap_content and I have keyword 'Michigan' in keywords, so it would switch it to the hyperlink that attached to keyword. Here is part of the function:

   def execute
      all_keys = Keyword.all.pluck(:key, :link).to_h.transform_keys(&:downcase)
      @post.swap_content = @post.swap_content.to_s.gsub!(/\w+/) do |word|
        url = all_keys[word.downcase]
        url ? "<a href='#{url}'>#{word}</a>" : word
      end
      @post.save!
  end

And my question is - how can I make it gsub only the first two keywords in @post.swap_content? For example, I have @post.swap_content 'michigan, michigan and michigan, utah and utah', how can I switch to hyperlinks only first two keywords(first two 'michigan' and first two 'utah')? I think, that I need somehow to work gsub but I don't know hot to manage number of words that can be gsub.

over 4 years ago · Santiago Trujillo
2 Respuestas
Responde la pregunta

0

You can provide a block to gsub that will be invoked with each match, you could use this to count occurences and condtionally replace content.

str = "Dog dog dog cat cat cat"
occurences = {}

str.gsub(/\w+/) do |match|
  # downcase so Dog and dog are counted together
  key = match.downcase
  # build a hash which counts the number of times we've matched a word.
  count = occurences.store(key, occurences.fetch(key, 0).next)
  
  # return the word unchanged or wrap in a hyperlink depending on count
  count > 2 ? match : "<a>#{match}</a>"
end

# output => "<a>Dog</a> <a>dog</a> dog <a>cat</a> <a>cat</a> cat"
over 4 years ago · Santiago Trujillo Denunciar

0

Suppose:

str = "Dog dog cat dog cat Dog cat cat"

If Ruby's regex engine supported variable-length negative lookbehinds we could write:

R = /\b(\w+)\b(?<!(?:\b\1\b.*){2})/i
str.gsub(R, '<a>\1</a>')
  #=> "<a>Dog</a> <a>dog</a> <a>cat</a> dog <a>cat</a> Dog cat cat"

We can write this regular expression in free-spacing mode to make it self-documenting:

R = /
    \b       # assert a word break
    (\w+)    # match 1+ word characters and save to capture group 1
    \b       # assert a word break
    (?!      # begin a negative lookbehind
      (?:    # begin a non-capture group
        \b   # assert a word break
        \1   # match the content of capture group 1
        \b   # assert a word break
        .*   # match 0+ characters
      )      # end non-capture group
      {2}    # execute non-capture group twice
    )        # end negative lookbehind
    /ix      # assert case-independent and free-spacing regex def modes

Unfortunately, Ruby's regex engine does not support variable-length (positive or negative) lookbehinds (though one day it might). It does, however, support variable-length (positive and negative) lookaheads. We therefore could reverse the string, perform the desired replacements using gsub then reverse the resulting string, as follows:

R = /\b(\w+)\b(?!(?:.*\b\1\b){2})/i
str.reverse.gsub(R, '>a/<\1>a<').reverse
  #=> "<a>Dog</a> <a>dog</a> <a>cat</a> dog <a>cat</a> Dog cat cat"

The steps are as follows.

s = str.reverse
  #=> "tac tac goD tac god tac god goD"
t = s.gsub(R, '>a/<\1>a<')
  #=> "tac tac goD >a/<tac>a< god >a/<tac>a< >a/<god>a< >a/<goD>a<"
t.reverse
  #=> "<a>Dog</a> <a>dog</a> <a>cat</a> dog <a>cat</a> Dog cat cat"

Let's have a closer look at the regular expression.

R = /
    \b       # assert a word break
    (\w+)    # match 1+ word characters and save to capture group 1
    \b       # assert a word break
    (?!      # begin a negative lookahead
      (?:    # begin a non-capture group
        .*   # match 0+ characters
        \b   # assert a word break
        \1   # match the content of capture group 1
        \b   # assert a word break
      )      # end non-capture group
      {2}    # execute non-capture group twice
    )        # end negative lookahead
    /ix      # assert case-independent and free-spacing regex def modes
over 4 years ago · Santiago Trujillo Denunciar
Responde la pregunta
Encuentra empleos remotos

¡Descubre la nueva forma de encontrar empleo!

Top de empleos
Top categorías de empleo
Empresas
Publicar vacante Precios Comercial
Legal
Términos y condiciones Política de privacidad
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Recomiéndame algunas ofertas
Necesito ayuda