Simple RAG in Rails: Answer Questions from Your Blog
Retrieve first, then generate
Retrieval-augmented generation gives a model relevant source text before it answers. For a Rails blog, start with approved articles and return citations to the original posts. The model is not a replacement for the database: Rails decides which content is visible and which passages are sent.
Build a small indexing pipeline
Extract plain text from Action Text, split it into passages of a few hundred words, and store each passage with its post ID and content digest. Generate embeddings with one chosen model, then store the vectors in PostgreSQL with pgvector. The vector column's dimensions must match that model's output. Use a background job, and remove or replace old passages when a post changes.
Keep retrieval behind one interface
class BlogAnswer
def initialize(retriever:, generator:)
@retriever = retriever
@generator = generator
end
def call(question:)
passages = @retriever.call(question: question, limit: 4)
return { answer: "No relevant approved article found.", sources: [] } if passages.empty?
{
answer: @generator.call(question: question, passages: passages),
sources: passages.map { |p| p.fetch(:post_id) }.uniq
}
end
end
The retriever and generator are application adapters, not built-in Rails APIs. The retriever embeds the question with the same embedding model, searches by vector distance, and joins the current posts table to require published and verified records. Apply a relevance cutoff calibrated against your own questions; nearest does not necessarily mean relevant.
Make the answer verifiable
Pass labeled excerpts to the generator with instructions to use only those excerpts and to admit missing evidence. Treat excerpts as data rather than commands. Render source links from the IDs Rails returned, and reject any generated citation outside that set. Escape model text before displaying it.
Check the failure cases
Try a question answered by one article, a question with no evidence, and a post withdrawn after indexing. Measure retrieval accuracy before tuning prose. Bound passage count and provider timeouts, avoid sending private drafts, and reindex all passages when changing embedding models.
Quy
Reactions
Comments
Sign in to join the conversation.