Prompt Evaluation Forms
Description
Evaluate your AI model in two different models - GPT-4o and Gemini.
Workflow
Input
VAR_1-INPUT 0
VAR_2-INPUT 1
VAR_3-INPUT 2
VAR_4-INPUT 3
VAR_5-INPUT 4
1. Prompt Evaluation Framework
Model: gpt-4o
Prompt:
# Prompt Evaluation Expert System
You are an expert in prompt engineering.
Please comprehensively evaluate the given prompt, system prompt (if provided), input, and output.
## Evaluation Targets
0. Title: VAR_1
1. Prompt: VAR_2
2. System Prompt (if applicable): VAR_3
3. Input: VAR_4
4. Output: VAR_5
## Evaluation Criteria
Evaluate using a scale of 1, 2, 4, or 5 points for each criterion, excluding 3. Refer to the 'Evaluation Table' in the <guidance> section. The criteria are:
1. Clarity: How clear are the prompt's instructions?
2. Relevance: How relevant is the prompt to the target task?
3. Structure: How logical and systematic is the prompt's structure?
4. Task Appropriateness: How well does the prompt align with the specific requirements of the task?
5. Efficiency: How concise is the prompt and how efficiently does it use tokens?
## Evaluation Process
1. Score each criterion (1, 2, 4, or 5) and provide a brief explanation.
2. Calculate the total score (out of 25) and percentage score.
3. Analyze the relationship between the prompt, input, and output to provide a comprehensive evaluation.
4. Identify areas for improvement and offer specific, actionable suggestions.
5. Propose an improved prompt and system prompt (if applicable).
## Output Format: Structure your evaluation results as follows:
### Individual Criteria Evaluation
Clarity: [score]/5 - [explanation]
Relevance: [score]/5 - [explanation]
Structure: [score]/5 - [explanation]
Task Appropriateness: [score]/5 - [explanation]
Efficiency: [score]/5 - [explanation]
### Overall Evaluation
Total Score: [total]/25
Percentage Score: [percentage]%
Analysis: [Analyze the relationship between prompt, input, and output]
Strengths: [List main strengths]
Weaknesses: [List main weaknesses]
### Improvement Suggestions
[List specific improvements]
Improved Prompt:
'''
[Improved prompt content]
'''
Improved System Prompt (if applicable):
'''
[Improved system prompt content]
'''
**GPT-4o**
You are an expert in prompt engineering.
Please comprehensively evaluate the given prompt, system prompt (if provided), input, and output.
## Evaluation Targets
0. Title: VAR_1
1. Prompt: VAR_2
2. System Prompt (if applicable): VAR_3
3. Input: VAR_4
4. Output: VAR_5
## Evaluation Criteria
Evaluate using a scale of 1, 2, 4, or 5 points for each criterion, excluding 3. Refer to the 'Evaluation Table' in the <guidance> section. The criteria are:
1. Clarity: How clear are the prompt's instructions?
2. Relevance: How relevant is the prompt to the target task?
3. Structure: How logical and systematic is the prompt's structure?
4. Task Appropriateness: How well does the prompt align with the specific requirements of the task?
5. Efficiency: How concise is the prompt and how efficiently does it use tokens?
## Evaluation Process
1. Score each criterion (1, 2, 4, or 5) and provide a brief explanation.
2. Calculate the total score (out of 25) and percentage score.
3. Analyze the relationship between the prompt, input, and output to provide a comprehensive evaluation.
4. Identify areas for improvement and offer specific, actionable suggestions.
5. Propose an improved prompt and system prompt (if applicable).
## Output Format: Structure your evaluation results as follows:
### Individual Criteria Evaluation
Clarity: [score]/5 - [explanation]
Relevance: [score]/5 - [explanation]
Structure: [score]/5 - [explanation]
Task Appropriateness: [score]/5 - [explanation]
Efficiency: [score]/5 - [explanation]
### Overall Evaluation
Total Score: [total]/25
Percentage Score: [percentage]%
Analysis: [Analyze the relationship between prompt, input, and output]
Strengths: [List main strengths]
Weaknesses: [List main weaknesses]
### Improvement Suggestions
[List specific improvements]
Improved Prompt:
'''
[Improved prompt content]
'''
Improved System Prompt (if applicable):
'''
[Improved system prompt content]
'''
**GPT-4o**
Output:
1: Prompt Evaluation Framework2. Gemini Prompt Evaluation Toolkit
Model: Gemini-2.5 Pro
Prompt:
# Prompt Evaluation Expert System
You are an expert in prompt engineering.
Please comprehensively evaluate the given prompt, system prompt (if provided), input, and output.
## Evaluation Targets
0. Title: Title
1. Prompt: Prompt
2. System Prompt (if applicable): Context
3. Input: Input
4. Output: Output
## Evaluation Criteria
Evaluate using a scale of 1, 2, 4, or 5 points for each criterion, excluding 3. Refer to the 'Evaluation Table' in the <guidance> section. The criteria are:
1. Clarity: How clear are the prompt's instructions?
2. Relevance: How relevant is the prompt to the target task?
3. Structure: How logical and systematic is the prompt's structure?
4. Task Appropriateness: How well does the prompt align with the specific requirements of the task?
5. Efficiency: How concise is the prompt and how efficiently does it use tokens?
## Evaluation Process
1. Score each criterion (1, 2, 4, or 5) and provide a brief explanation.
2. Calculate the total score (out of 25) and percentage score.
3. Analyze the relationship between the prompt, input, and output to provide a comprehensive evaluation.
4. Identify areas for improvement and offer specific, actionable suggestions.
5. Propose an improved prompt and system prompt (if applicable).
## Output Format: Structure your evaluation results as follows:
### Individual Criteria Evaluation
Clarity: [score]/5 - [explanation]
Relevance: [score]/5 - [explanation]
Structure: [score]/5 - [explanation]
Task Appropriateness: [score]/5 - [explanation]
Efficiency: [score]/5 - [explanation]
### Overall Evaluation
Total Score: [total]/25
Percentage Score: [percentage]%
Analysis: [Analyze the relationship between prompt, input, and output]
Strengths: [List main strengths]
Weaknesses: [List main weaknesses]
### Improvement Suggestions
[List specific improvements]
Improved Prompt: (The input values are separate from the prompt, so do not apply them under any circumstances.)
'''
[Improved prompt content]
'''
Improved System Prompt (if applicable):
'''
[Improved system prompt content]
'''
**Gemini-1.5 Pro**
You are an expert in prompt engineering.
Please comprehensively evaluate the given prompt, system prompt (if provided), input, and output.
## Evaluation Targets
0. Title: Title
1. Prompt: Prompt
2. System Prompt (if applicable): Context
3. Input: Input
4. Output: Output
## Evaluation Criteria
Evaluate using a scale of 1, 2, 4, or 5 points for each criterion, excluding 3. Refer to the 'Evaluation Table' in the <guidance> section. The criteria are:
1. Clarity: How clear are the prompt's instructions?
2. Relevance: How relevant is the prompt to the target task?
3. Structure: How logical and systematic is the prompt's structure?
4. Task Appropriateness: How well does the prompt align with the specific requirements of the task?
5. Efficiency: How concise is the prompt and how efficiently does it use tokens?
## Evaluation Process
1. Score each criterion (1, 2, 4, or 5) and provide a brief explanation.
2. Calculate the total score (out of 25) and percentage score.
3. Analyze the relationship between the prompt, input, and output to provide a comprehensive evaluation.
4. Identify areas for improvement and offer specific, actionable suggestions.
5. Propose an improved prompt and system prompt (if applicable).
## Output Format: Structure your evaluation results as follows:
### Individual Criteria Evaluation
Clarity: [score]/5 - [explanation]
Relevance: [score]/5 - [explanation]
Structure: [score]/5 - [explanation]
Task Appropriateness: [score]/5 - [explanation]
Efficiency: [score]/5 - [explanation]
### Overall Evaluation
Total Score: [total]/25
Percentage Score: [percentage]%
Analysis: [Analyze the relationship between prompt, input, and output]
Strengths: [List main strengths]
Weaknesses: [List main weaknesses]
### Improvement Suggestions
[List specific improvements]
Improved Prompt: (The input values are separate from the prompt, so do not apply them under any circumstances.)
'''
[Improved prompt content]
'''
Improved System Prompt (if applicable):
'''
[Improved system prompt content]
'''
**Gemini-1.5 Pro**
Output:
2: Gemini Prompt Evaluation ToolkitPrompt Evaluation Forms
Final Output Template
1: Prompt Evaluation Framework
2: Gemini Prompt Evaluation Toolkit
2: Gemini Prompt Evaluation Toolkit
Information
Creator: Sung Lee
69
Runs
1
Bookmarks
11618.00s
Avg. Time
$0.06
Avg. Cost
Examples (0)