試験CCAR-P トピック1 問題28 スレッド
Anthropic CCAR-Pのリアル試験問題集
問題 #: 28
トピック #: 1
問題 #: 28
トピック #: 1
You are running a controlled experiment to compare two prompts and must complete the design steps before executing the experiment.
Which two steps must be completed BEFORE running the experiment with random assignment? (Select two.) Each correct answer presents part of the solution.
Which two steps must be completed BEFORE running the experiment with random assignment? (Select two.) Each correct answer presents part of the solution.
おすすめの解答:A,C 解答を投票する
A controlled prompt experiment must begin with a falsifiable hypothesis and a predefined primary metric.
Option C prevents the team from examining results first and then selecting whichever metric makes the candidate look successful. The metric might measure task accuracy, rubric score, citation validity, escalation rate, latency, cost, or another criterion directly connected to the hypothesis.
Option A determines whether the experiment can detect a practically meaningful improvement. The minimum detectable effect expresses the smallest difference worth acting upon, while the power calculation determines the required sample size. Without this step, the experiment may be too small to detect a real improvement or unnecessarily large and expensive.
Random assignment should then distribute representative traffic between the control and candidate prompts while controlling model version, retrieval configuration, tool availability, and other confounding variables.
Options B and D occur after data collection. Option E follows the completed analysis and decision. The team should also define significance thresholds, stopping rules, guardrail metrics, exclusion criteria, and treatment of repeated observations before launch.
Study Guide references/topics: Prompt A/B testing; hypothesis definition; primary metrics; minimum detectable effect; statistical power; random assignment; decision sequencing.
Option C prevents the team from examining results first and then selecting whichever metric makes the candidate look successful. The metric might measure task accuracy, rubric score, citation validity, escalation rate, latency, cost, or another criterion directly connected to the hypothesis.
Option A determines whether the experiment can detect a practically meaningful improvement. The minimum detectable effect expresses the smallest difference worth acting upon, while the power calculation determines the required sample size. Without this step, the experiment may be too small to detect a real improvement or unnecessarily large and expensive.
Random assignment should then distribute representative traffic between the control and candidate prompts while controlling model version, retrieval configuration, tool availability, and other confounding variables.
Options B and D occur after data collection. Option E follows the completed analysis and decision. The team should also define significance thresholds, stopping rules, guardrail metrics, exclusion criteria, and treatment of repeated observations before launch.
Study Guide references/topics: Prompt A/B testing; hypothesis definition; primary metrics; minimum detectable effect; statistical power; random assignment; decision sequencing.
小杉** 2026-08-28 01:43:29
コメント
他人の解答コメントを賛成するのも、その解答に一票を入れることになります。したがって、すでに同じ意見の投票コメントが存在する場合、新規コメントをする代わりに賛成することもできます。
コメントを通報する
コメント中
今すぐ 新規登録 / ログイン (無料です)。