Skip to main content
Rails Engineering

Publish, review, and manage your writing.

#rails #ruby #elasticsearch

Ruby and Rails Deep Dive: Designing a Consistent Search Pipeline

M
Manh
• • 5 mins read
Technical article hero

Define what readers may find

A public search page should expose the same records as the public feed. For a moderated blog, that means published and verified posts. Put these restrictions in the search query, rather than filtering a page of results afterward and producing misleading totals.

query = params[:q].to_s.strip.first(200)
@posts = Post.search(
  query.presence || "*",
  fields: ["title^3", "body", "tags"],
  where: { status: Post.statuses.fetch("published"), verified: true },
  page: params[:page], per_page: 20
)

This Searchkick example assumes title, body, tags, status, and verified are present in the indexed document. The title boost favors direct matches, but relevance should be checked against real reader questions rather than tuned by intuition alone.

Keep the index consistent

Index creation, body edits, tag changes, publication, and unpublication. Rich text and join tables can change without a normal title update, so check those paths explicitly. Queue indexing after a successful commit. Treat the database as the source of truth and make a full rebuild an explicit maintenance operation.

Handle an unavailable search service

Show a clear temporary failure message, or use a bounded database fallback with the same visibility restrictions. Do not silently return every post when the search engine fails. Keep result URLs and titles usable with keyboard navigation.

Evaluate quality

Create a small list of expected queries: exact titles, common Rails terms, misspellings, and phrases that should return nothing. Check both ranking and authorization. Also test an approved post that becomes private while the index is stale; recheck visibility before rendering sensitive records.

Design a Ruby boundary you can reason about

Keep HTTP parameters outside the indexing layer. Pass explicit keyword arguments to a small service and inject its external dependencies. Ruby's duck typing lets a production client and a deterministic test fake share the same call contract without inheriting from a common framework class. Document what that contract returns and which failures it raises; implicit behavior is harder to operate than a small explicit interface.

class SearchDocument
  def self.build(post)
    {
      id: post.id,
      title: post.title.to_s.dup.freeze,
      body: post.body.to_plain_text.freeze,
      tags: post.tags.map { |tag| tag.name.dup.freeze }.sort.freeze,
      visible: post.published? && post.verified?
    }.freeze
  end
end

Freezing a hash is shallow: its strings and arrays remain mutable unless you freeze them too. This snapshot freezes the nested values it constructs, and duplicates model-backed strings before freezing them. It is an example payload builder, not a replacement for this application's Searchkick schema. Avoid memoizing it across requests: a cached document can outlive an approval change.

Know when Ruby executes SQL

An Active Record relation describes a query until an operation materializes it. Iteration loads model objects; pluck asks the database for selected values; converting a relation into an array shifts subsequent filtering into Ruby. For an indexing batch, preload the author and tags, then extract rich text deliberately. A low SQL count does not guarantee low memory use when each record contains a long article.

Measure allocations and batch memory alongside database time. Avoid loading an entire corpus into one array. Choose a batch size using realistic body lengths, and bound the payload size sent to the search engine. A worker that is fast for a hundred short posts may fail on a thousand large ones.

Serialize conflicting publication changes

Imagine one administrator approves a post while another withdraws it. Reading state, deciding, and saving in separate steps allows both decisions to use stale data. A row lock can serialize the critical section. Acquire it before changing attributes, keep it short, and do not make network requests while holding it.

post.with_lock do
  raise "Admin required" unless actor.admin?
  post.verify!(actor)
end

This example uses this blog's verify! method, which rejects draft posts. with_lock reloads the persisted record under a database lock and wraps the block in a transaction. All competing state transitions must follow the same locking discipline. For a human edit form, optimistic locking with a lock_version column may instead be preferable because it can report a conflict rather than silently replacing another edit.

Close the commit-to-queue failure window

An after_commit callback avoids indexing a transaction that later rolls back. It does not make a database commit and a Redis enqueue atomic: the process can stop between them. If search consistency matters, write an outbox event in the same database transaction as the state change. A separate dispatcher reads pending events and delivers them with retries.

The outbox needs an event ID, post ID, revision or content digest, delivery state, and a unique constraint for the chosen event identity. Claim work safely when several dispatchers run. Mark delivery only after the destination acknowledges it. A crash after delivery but before marking completion produces a duplicate, so the consumer still has to tolerate repeated events. These tables and workers are design extensions, not components already supplied by Rails.

Handle stale and reordered jobs

A delayed job may contain an old approved snapshot after a newer withdrawal. Prefer jobs carrying a post ID that rebuild from current committed state. Two workers can still read different revisions and finish in the opposite order, so use per-post serialization or destination-side version checks when strict ordering matters. Define deletion behavior explicitly: a missing source record should remove its indexed document rather than fail forever.

Rich text, tags, and approval all contribute to the document. A fingerprint over their canonical values can identify meaningful content changes, but a digest alone does not provide ordering. Keep a monotonic revision if the consumer needs to reject older writes. Reconcile database records and index documents periodically to repair missed events.

Make failure observable and testable

Record event ID, post ID, revision, attempt count, and latency. Monitor the age of the oldest pending event: workers can be healthy while the backlog grows. Use bounded retries with jitter for transient failures and a visible failed-event queue for malformed payloads. Keep article text and credentials out of routine error logs.

Test rollback without an event, duplicate delivery, withdrawal during indexing, and older writes arriving last. Exercise concurrency with separate database connections and synchronization barriers rather than arbitrary sleeps. Inject a client that fails before acknowledgment and one that succeeds before the dispatcher crashes. Verify the final document and public visibility, not merely that a method was called.

Choose complexity from the requirement

A small internal blog may accept a short indexing delay with after-commit jobs and a repair task. A large publication workflow may justify an outbox and versioned writes. State the tolerated delay, recovery process, and privacy boundary before selecting the design. Always recheck current visibility when serving sensitive results; an eventually consistent index should not decide access rights.

References: Active Record queries, transaction callbacks, and pessimistic locking.

Share this article
Author

Manh

Reactions

React Hover to preview

Comments

Sign in to join the conversation.

Recommended Reading