AI features
プロンプト変更による回帰を防ぐ評価ハーネス

CIで実行され、プロンプトやモデルの変更が状況を悪化させたときに通知するAI機能のテストスイート。
What you get
- 実際のプロダクショントラフィックから構築されたゴールデンデータセット
- アサーションとモデル採点ルーブリックを組み合わせた自動採点
- スコアがベースラインを下回った場合にビルドを失敗させるCIジョブ
- モデルまたはプロンプトバージョン切り替えのための比較レポート
Usually built with
Python
Promptfoo
GitHub Actions
OpenAI API
A guide, not a requirement — sellers propose what suits your situation.
Nobody offers this yet
No seller has published this service yet. Ask anyway — the request goes to the marketplace and sellers who do this kind of work can answer it.
Or publish it yourself if this is your work.
Close to this
Make your RAG bot stop inventing answersDiagnose why retrieval misses, then fix chunking, reranking and prompting until answers are grounded and citable.from €2,200MCP server so agents can use your internal toolsExpose your internal APIs to coding and support agents safely, with scoped permissions and a full audit trail.from €2,000Streaming chat UI with tool calls and real error statesThe front end your AI feature deserves: token streaming, cancellation, retries and messages that survive a refresh.from €2,000Moderation for user-generated text and imagesAutomatic screening of uploads and posts with tunable thresholds, an appeals path and a queue for the grey area.from €2,400Meeting transcription with searchable summariesCalls transcribed with speaker labels, summarised into decisions and actions, and searchable across every past meeting.from €2,500Harden your LLM feature against prompt injectionClose the gaps that let a crafted input make your agent leak data, call the wrong tool or ignore its instructions.from €1,800
What will it cost?
Four questions, an instant range. No account, no waiting.
Estimated range
€2,200 – €4,200
6-10
Based on the catalogue guide — nobody has published this service yet, so there is no market price to work from. An estimate, not a quote: a seller prices the real job once they have read it.
Request a quote
Nobody has published this service yet.