Index

November 24, 2024
in RAG
4 min read

RAG is a fancy way of stuffing additional information into the prompt of a language model. By giving the model more information, we can get more contextual responses that are contextually relevant to what we need. But don't language models already have access to all of the world's information?

Imagine you're starting a new job. Would you rather:

Have access to all of Wikipedia and hope the information you need is somewhere in there
Have your company's specific documentation, procedures, and guidelines

November 23, 2024
in LLMs, Synthetic Data
6 min read

Synthetic Data is no Free Lunch

I spent some time playing with a new framework called Dria recently that uses LLMs to generate synthetic data. I couldn't get it to work but I did spend some time digging through their source code, and I thought I'd share some of my thoughts on the topic.

Over the past few weeks, I've generated a few million tokens of synthetic data for some projects. I'm still figuring out the best way to do it but I think it's definitely taught me that it's no free lunch. You do need to spend some time thinking about how to generate the data that you want.

The Premise

An example

When I first started generating synthetic data for question-answering systems, I thought it would be straightforward - all I had to do was to ask a language model to generate a few thousand questions that a user might ask.

November 20, 2024
in LLMs, Applied AI
4 min read

You're probably not doing experiments right

I recently started working as a research engineer and it's been a significant mindset shift in how I approach my work. it's tricky to run experiments with LLMs efficiently and accurately and after months of trial and error, I've found that there are three key factors that make the biggest difference

Being clear about what you're varying
Investing time to build out some infrastructure
Doing some simple sensitivity analysis

Let's see how each of these can make a difference in your experimental workflow.

September 21, 2024
in LLMs, langchain, Instructor
7 min read

Why Instructor might be a better bet than Langchain

Introduction

If you're building LLM applications, a common question is which framework to use: Langchain, Instructor, or something else entirely. I've found that this decision really comes down to a few critical factors to choose the right one for your application. We'll do so in three parts

First we'll talk about testing and granular controls and why you should be thinking about it from the start
Then we'll explain why you should be evaluating a framework's ability to experiment quickly with different models and prompts and adopt new features quickly.
Finally, we'll consider why long term maintenance is also an important factor and why Instructor often provides a balanced solution, offering both simplicity and flexibility.

September 8, 2024
in Instructor
8 min read

How does Instructor work?

For Python developers working with large language models (LLMs), instructor has become a popular tool for structured data extraction. While its capabilities may seem complex, the underlying mechanism is surprisingly straightforward. In this article, we'll walk through a high level overview of how the library works and how we support the OpenAI Client.

We'll start by looking at

Why should you care about Structured Extraction?
What is the high level flow
How does a request go from Pydantic Model to Validated Function Call?

By the end of this article, you'll have a good understand of how instructor helps you get validated outputs from your LLM calls and a better understanding of how you might be able to contribute to the library yourself.

September 5, 2024
in Evals, Braintrust
8 min read

Getting Started with Evals - a speedrun through Braintrust

For software engineers struggling with LLM application performance, simple evaluations are your secret weapon. Forget the complexity — we'll show you how to start testing your LLM in just 5 minutes using Braintrust. By the end of this article, you'll have a working example of a test harness that you can easily customise for your own use cases.

We'll be using a cleaned version of the GSM8k dataset that you can find here.

Here's what we'll cover:

Setting up Braintrust
Writing our first task to evaluate an LLM's response to the GSM8k with Instructor
Simple recipes that you'll need

August 27, 2024
in LLMs, Synthetic Data
5 min read

How to create synthetic data that works

Synthetic data can accelerate AI development, but generating high-quality datasets remains challenging. In this article, I'll walk through a few experiments I've done with synthetic data generation and the takeaways I've learnt so that you can do the same.

We'll do by covering

Limitations of simple generation methods : Why simple generation methods produce homogeneous data
Entropy and why it matters : Techniques to increase diversity in synthetic datasets
Practical Implementations : Some simple examples of how to increase entropy and diversity to get better synthetic data

June 30, 2024
in AI Engineering, LLMs
5 min read

AI Engineering World Fair

What's new?

Last year, we saw a lot of interest in the use of LLMs for new use cases. This year, with more funding and interest in the space, we've finally started thinking about productionizing these models at scale and making sure that they're reliable, consistent and secure.

Let's start with a few definitions

Agent : This is a LLM which is provided with a few tools it can call. The agentic part of this system comes from the ability to make decisions based on some input. This is similar to Harrison Chase's article here
Evaluations : A set of metrics that we can look at to understand where our current system falls short. An example could be measuring precision and recall.
Synthethic Data Generation: Data generated by a LLM which is meant to mimic real data

May 2, 2024
in LLMs, Walkthrough
19 min read

Grokking LLMs

I've spent the last year working with LLMs and writing a good amount of technical content on how to use them effectively, mostly with the help of structured parsing using a framework like Instructor. Most of what I know now is self-taught and this is the guide that I wish I had when starting out.

It should take about 10-15 minutes at most to read and I've added some resources along the way that are relevant to you. If you're looking for a higher level, i suggest skimming over the first two sections and then focusing more on the application/data side of things!

I hope that after reading this essay, you walk away with an enthusiasm that these models are going to change so much things that we know today. We have models with reasoning abilities and knowledge capacities that dwarf many humans today in tasks such as Mathetical Reasoning, QnA and more.

April 27, 2024
6 min read

Introduction

It's really fun to create your own tools. With some extra time on my hands this weekend, I decided to work on building a small tool that would solve a problem i'd been facing for some time - converting wikilinks to relative links.

For those who are unaware, when you work in tools like Obsidian, the default tends to be wikilinks that look like this [[wiki-link]]. This is great if you're only using obsidian but limits the portability of your markdown script itself. For platforms such as Github, the lack of absolute links means that you can't easily click and navigate between markdown files on their web platform.

Index

Is RAG dead?

What is RAG?

Synthetic Data is no Free Lunch

The Premise

An example

You're probably not doing experiments right

Why Instructor might be a better bet than Langchain

Introduction

How does Instructor work?

Getting Started with Evals - a speedrun through Braintrust

How to create synthetic data that works

AI Engineering World Fair

What's new?

Grokking LLMs

Introduction