Business
Jobs
  • About Us
  • Solutions
    • Job Postings
      Post your job and receive qualified candidates in 48h.
    • Candidate Assessments
      500+ technical and psychological tests, plus anti-fraud.
    • Headhunting
      Tailor-made executive search from start to finish.
    • Payroll + EOR
      Payroll dispersal and EOR across 15+ LATAM countries.
  • Pricing
  • Jobs

0

333
Views
Ruby gsub for exact number of words

I have a piece of code, where I can switch words from @post.swap_content to hyperlinks by keyword. For example, I have a word 'michigan' in @post.swap_content and I have keyword 'Michigan' in keywords, so it would switch it to the hyperlink that attached to keyword. Here is part of the function:

   def execute
      all_keys = Keyword.all.pluck(:key, :link).to_h.transform_keys(&:downcase)
      @post.swap_content = @post.swap_content.to_s.gsub!(/\w+/) do |word|
        url = all_keys[word.downcase]
        url ? "<a href='#{url}'>#{word}</a>" : word
      end
      @post.save!
  end

And my question is - how can I make it gsub only the first two keywords in @post.swap_content? For example, I have @post.swap_content 'michigan, michigan and michigan, utah and utah', how can I switch to hyperlinks only first two keywords(first two 'michigan' and first two 'utah')? I think, that I need somehow to work gsub but I don't know hot to manage number of words that can be gsub.

over 4 years ago · Santiago Trujillo
2 answers
Answer question

0

You can provide a block to gsub that will be invoked with each match, you could use this to count occurences and condtionally replace content.

str = "Dog dog dog cat cat cat"
occurences = {}

str.gsub(/\w+/) do |match|
  # downcase so Dog and dog are counted together
  key = match.downcase
  # build a hash which counts the number of times we've matched a word.
  count = occurences.store(key, occurences.fetch(key, 0).next)
  
  # return the word unchanged or wrap in a hyperlink depending on count
  count > 2 ? match : "<a>#{match}</a>"
end

# output => "<a>Dog</a> <a>dog</a> dog <a>cat</a> <a>cat</a> cat"
over 4 years ago · Santiago Trujillo Report

0

Suppose:

str = "Dog dog cat dog cat Dog cat cat"

If Ruby's regex engine supported variable-length negative lookbehinds we could write:

R = /\b(\w+)\b(?<!(?:\b\1\b.*){2})/i
str.gsub(R, '<a>\1</a>')
  #=> "<a>Dog</a> <a>dog</a> <a>cat</a> dog <a>cat</a> Dog cat cat"

We can write this regular expression in free-spacing mode to make it self-documenting:

R = /
    \b       # assert a word break
    (\w+)    # match 1+ word characters and save to capture group 1
    \b       # assert a word break
    (?!      # begin a negative lookbehind
      (?:    # begin a non-capture group
        \b   # assert a word break
        \1   # match the content of capture group 1
        \b   # assert a word break
        .*   # match 0+ characters
      )      # end non-capture group
      {2}    # execute non-capture group twice
    )        # end negative lookbehind
    /ix      # assert case-independent and free-spacing regex def modes

Unfortunately, Ruby's regex engine does not support variable-length (positive or negative) lookbehinds (though one day it might). It does, however, support variable-length (positive and negative) lookaheads. We therefore could reverse the string, perform the desired replacements using gsub then reverse the resulting string, as follows:

R = /\b(\w+)\b(?!(?:.*\b\1\b){2})/i
str.reverse.gsub(R, '>a/<\1>a<').reverse
  #=> "<a>Dog</a> <a>dog</a> <a>cat</a> dog <a>cat</a> Dog cat cat"

The steps are as follows.

s = str.reverse
  #=> "tac tac goD tac god tac god goD"
t = s.gsub(R, '>a/<\1>a<')
  #=> "tac tac goD >a/<tac>a< god >a/<tac>a< >a/<god>a< >a/<goD>a<"
t.reverse
  #=> "<a>Dog</a> <a>dog</a> <a>cat</a> dog <a>cat</a> Dog cat cat"

Let's have a closer look at the regular expression.

R = /
    \b       # assert a word break
    (\w+)    # match 1+ word characters and save to capture group 1
    \b       # assert a word break
    (?!      # begin a negative lookahead
      (?:    # begin a non-capture group
        .*   # match 0+ characters
        \b   # assert a word break
        \1   # match the content of capture group 1
        \b   # assert a word break
      )      # end non-capture group
      {2}    # execute non-capture group twice
    )        # end negative lookahead
    /ix      # assert case-independent and free-spacing regex def modes
over 4 years ago · Santiago Trujillo Report
Answer question
Find remote jobs

Discover the new way to find a job!

Top jobs
Top job categories
Business
Post vacancy Pricing Sales
Legal
Terms and conditions Privacy policy
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Show me some job opportunities
There's an error!