Skip to main content

1. Installing FutureAGI

2. Loading Dataset

Dataset used here contains value proposition and Linkedin posts, using which the AI models will create openers as per the prompts.
Below the sample of dataset used in this cookbook:

3. Initialising Future AGI’s Evaluator Client

4. Defining Custom Deterministic Eval

Definig custom deterministic eval that is tailored to our use case. With below config:
Click here to learn how to create your own custom eval.

5. Defining Judging Criteria for Evaluating AI Generated Openers

  • Here, the AI generated opener is being judged on following criteries:
    • Engagement
    • Tone
    • Relevance
    • Appropriateness
    • Impact
  • You can include more criterias that suits your use-case, given that you explicitily define on how to choose the tags.
  • Tags are nothing but the output deterministic eval returns. Depending on the use-case, you can choose multi-choice or single-choice.
  • You can add any number of tags, given that you have defined on how to choose those tags.

6. Evaluating AI Generated Openers Using Custom Deterministic Eval

Below code will create test case for each judging criteria using custom deterministic eval. Since we are using f-string in the “opener” and it requires input keys inside double curly braces, so include them inside 4 curly braces. Otherwise the eval would not recieve these inputs to perform correct evaluation.
Output:

Evaluation on Prompt 1

Engagement Evaluation Result 1Tone Evaluation Result 1Relevance Evaluation Result 1Appropriateness Evaluation Result 1Impact Evaluation Result 1
GoodGoodGoodPoorGood
GoodGoodGoodGoodGood
GoodGoodGoodPoorGood
GoodGoodGoodGoodGood

Evaluation on Prompt 2

Engagement Evaluation Result 2Tone Evaluation Result 2Relevance Evaluation Result 2Appropriateness Evaluation Result 2Impact Evaluation Result 2
PoorGoodGoodGoodGood
GoodGoodGoodGoodGood
GoodGoodPoorPoorPoor
GoodGoodGoodGoodGood

7. Selecting Winner Prompt

For our use-case, that prompt is considered as a winner prompt that performs better on these judging criterias. Performance of a prompt can be judged by taking the majority of positve tags, here “Good” across all column per row. Then both the prompts are compared, and whichever has more number of “Good” prompts will be considered as a winner prompt.
Output:
Eval Rating Prompt 1Eval Rating Prompt 2
GoodGood
GoodGood
GoodGood
GoodPoor
GoodGood
GoodGood
GoodGood
GoodGood
GoodGood
GoodGood
Winner Prompt: Prompt 1