chore(skore): Align default metrics with is_computable and discouraged - #3195
chore(skore): Align default metrics with is_computable and discouraged#3195glemaitre wants to merge 5 commits into
Conversation
Co-authored-by: Guillaume Lemaitre <[email protected]>
… inserting it Co-authored-by: Guillaume Lemaitre <[email protected]>
Coverage Report for |
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| File | Stmts | Miss | Branch | BrPart | Cover | Missing |
|---|---|---|---|---|---|---|
| skore/src/skore | ||||||
| __init__.py | 41 | 2 | 2 | 0 | 95% | 118–119 |
| _config.py | 58 | 3 | 12 | 1 | 94% | 70, 117–118 |
| exceptions.py | 4 | 4 | 0 | 0 | 0% | 4, 15, 19, 23 |
| skore/src/skore/_plugins | ||||||
| __init__.py | 12 | 0 | 0 | 0 | 100% | |
| serde.py | 32 | 0 | 8 | 0 | 100% | |
| skore/src/skore/_plugins/hub | ||||||
| __init__.py | 9 | 2 | 0 | 0 | 77% | 15, 20 |
| exception.py | 2 | 0 | 0 | 0 | 100% | |
| json.py | 10 | 1 | 2 | 1 | 90% | 16 |
| metric.py | 34 | 2 | 6 | 2 | 94% | 74, 82 |
| skore/src/skore/_plugins/hub/artifact | ||||||
| __init__.py | 0 | 0 | 0 | 0 | 100% | |
| artifact.py | 23 | 0 | 4 | 0 | 100% | |
| serializer.py | 27 | 0 | 2 | 0 | 100% | |
| upload.py | 26 | 0 | 4 | 0 | 100% | |
| skore/src/skore/_plugins/hub/artifact/media | ||||||
| __init__.py | 5 | 0 | 0 | 0 | 100% | |
| data.py | 23 | 0 | 0 | 0 | 100% | |
| inspection.py | 58 | 0 | 10 | 0 | 100% | |
| media.py | 12 | 0 | 0 | 0 | 100% | |
| model.py | 10 | 0 | 0 | 0 | 100% | |
| performance.py | 57 | 0 | 0 | 0 | 100% | |
| skore/src/skore/_plugins/hub/artifact/pickle | ||||||
| __init__.py | 2 | 0 | 0 | 0 | 100% | |
| pickle.py | 24 | 0 | 2 | 0 | 100% | |
| skore/src/skore/_plugins/hub/authentication | ||||||
| __init__.py | 0 | 0 | 0 | 0 | 100% | |
| apikey.py | 7 | 0 | 0 | 0 | 100% | |
| login.py | 28 | 4 | 4 | 2 | 85% | 37, 42–43, 52 |
| token.py | 80 | 0 | 8 | 0 | 100% | |
| uri.py | 6 | 0 | 0 | 0 | 100% | |
| skore/src/skore/_plugins/hub/client | ||||||
| __init__.py | 0 | 0 | 0 | 0 | 100% | |
| client.py | 88 | 10 | 18 | 3 | 88% | 140, 187–189, 191–192, 194, 196, 198, 230 |
| skore/src/skore/_plugins/hub/project | ||||||
| __init__.py | 0 | 0 | 0 | 0 | 100% | |
| project.py | 143 | 6 | 28 | 5 | 95% | 88, 113, 128, 327, 425, 455 |
| skore/src/skore/_plugins/hub/report | ||||||
| __init__.py | 3 | 0 | 0 | 0 | 100% | |
| cross_validation_report.py | 185 | 6 | 38 | 3 | 96% | 185–189, 204 |
| estimator_report.py | 17 | 0 | 0 | 0 | 100% | |
| report.py | 56 | 0 | 4 | 0 | 100% | |
| skore/src/skore/_plugins/local | ||||||
| __init__.py | 2 | 0 | 0 | 0 | 100% | |
| project.py | 250 | 3 | 58 | 2 | 98% | 455, 497, 518 |
| skore/src/skore/_plugins/mlflow | ||||||
| __init__.py | 5 | 0 | 0 | 0 | 100% | |
| project.py | 281 | 23 | 72 | 10 | 91% | 47, 49–51, 129, 236, 254, 272–273, 362, 413, 521, 523, 539, 541–546, 548–550 |
| reports.py | 160 | 6 | 36 | 5 | 96% | 129, 176, 214–215, 277, 288 |
| skore/src/skore/_project | ||||||
| __init__.py | 0 | 0 | 0 | 0 | 100% | |
| _summary.py | 118 | 1 | 48 | 1 | 99% | 108 |
| dependencies.py | 19 | 0 | 6 | 0 | 100% | |
| git.py | 30 | 0 | 4 | 0 | 100% | |
| login.py | 17 | 3 | 6 | 2 | 82% | 57, 66–67 |
| plugin.py | 9 | 0 | 0 | 0 | 100% | |
| project.py | 65 | 1 | 18 | 1 | 98% | 151 |
| protocol.py | 5 | 5 | 0 | 0 | 0% | 3, 5, 7, 13, 16 |
| types.py | 4 | 0 | 0 | 0 | 100% | |
| skore/src/skore/_sklearn | ||||||
| __init__.py | 8 | 0 | 0 | 0 | 100% | |
| _base.py | 158 | 6 | 42 | 6 | 96% | 71, 222, 299, 308, 314, 349 |
| compare.py | 5 | 0 | 0 | 0 | 100% | |
| evaluate.py | 44 | 0 | 22 | 0 | 100% | |
| feature_names.py | 28 | 0 | 12 | 0 | 100% | |
| find_ml_task.py | 61 | 0 | 46 | 1 | 100% | |
| metrics.py | 404 | 1 | 104 | 1 | 99% | 484 |
| train_test_split.py | 17 | 0 | 0 | 0 | 100% | |
| types.py | 20 | 1 | 0 | 0 | 95% | 31 |
| skore/src/skore/_sklearn/_checks | ||||||
| __init__.py | 3 | 0 | 0 | 0 | 100% | |
| _utils.py | 126 | 0 | 54 | 1 | 100% | |
| accessor.py | 35 | 0 | 14 | 0 | 100% | |
| base.py | 104 | 0 | 28 | 2 | 100% | |
| model_checks.py | 468 | 1 | 152 | 2 | 99% | 896 |
| tunable_hyperparameters.py | 5 | 0 | 0 | 0 | 100% | |
| skore/src/skore/_sklearn/_comparison | ||||||
| __init__.py | 9 | 0 | 0 | 0 | 100% | |
| inspection_accessor.py | 25 | 0 | 2 | 0 | 100% | |
| metrics_accessor.py | 116 | 2 | 26 | 2 | 98% | 433, 688 |
| report.py | 191 | 7 | 90 | 6 | 96% | 276–277, 481, 600, 682–684 |
| skore/src/skore/_sklearn/_cross_validation | ||||||
| __init__.py | 11 | 0 | 0 | 0 | 100% | |
| data_accessor.py | 40 | 2 | 14 | 2 | 95% | 49, 75 |
| inspection_accessor.py | 25 | 0 | 2 | 0 | 100% | |
| metrics_accessor.py | 101 | 5 | 18 | 2 | 95% | 107–108, 110, 113, 637 |
| report.py | 205 | 9 | 46 | 5 | 95% | 230, 235, 246, 251, 353, 595, 736–738 |
| skore/src/skore/_sklearn/_estimator | ||||||
| __init__.py | 11 | 0 | 0 | 0 | 100% | |
| data_accessor.py | 61 | 1 | 28 | 1 | 98% | 65 |
| inspection_accessor.py | 42 | 1 | 8 | 2 | 97% | 298 |
| metrics_accessor.py | 116 | 0 | 22 | 0 | 100% | |
| report.py | 324 | 10 | 88 | 7 | 96% | 64, 80, 299, 381, 435, 685, 770, 865–867 |
| skore/src/skore/_sklearn/_plot | ||||||
| __init__.py | 3 | 0 | 0 | 0 | 100% | |
| base.py | 75 | 2 | 14 | 1 | 97% | 72–73 |
| utils.py | 160 | 3 | 76 | 4 | 98% | 254–255, 454 |
| skore/src/skore/_sklearn/_plot/data | ||||||
| __init__.py | 2 | 0 | 0 | 0 | 100% | |
| table_report.py | 189 | 1 | 60 | 2 | 99% | 727 |
| skore/src/skore/_sklearn/_plot/inspection | ||||||
| __init__.py | 0 | 0 | 0 | 0 | 100% | |
| calibration_curve.py | 82 | 11 | 14 | 8 | 86% | 120–123, 125, 130–131, 234, 237–238, 268 |
| coefficients.py | 202 | 0 | 98 | 1 | 100% | |
| impurity_decrease.py | 104 | 2 | 34 | 3 | 98% | 444, 488 |
| permutation_importance.py | 201 | 1 | 90 | 1 | 99% | 659 |
| utils.py | 32 | 0 | 10 | 0 | 100% | |
| skore/src/skore/_sklearn/_plot/metrics | ||||||
| __init__.py | 6 | 0 | 0 | 0 | 100% | |
| confusion_matrix.py | 200 | 0 | 66 | 2 | 100% | |
| metrics_summary_display.py | 176 | 0 | 62 | 1 | 100% | |
| precision_recall_curve.py | 118 | 0 | 32 | 1 | 100% | |
| prediction_error.py | 178 | 0 | 58 | 2 | 100% | |
| roc_curve.py | 122 | 0 | 34 | 2 | 100% | |
| skore/src/skore/_utils | ||||||
| __init__.py | 6 | 2 | 0 | 0 | 66% | 8, 13 |
| _accessor.py | 106 | 20 | 30 | 11 | 81% | 13, 36, 61–65, 68, 70–71, 76, 81, 83, 85, 92–94, 164, 216, 236 |
| _cache.py | 37 | 0 | 2 | 1 | 100% | |
| _cache_key.py | 38 | 6 | 24 | 6 | 84% | 22, 24, 26, 53, 61, 70 |
| _callable.py | 9 | 0 | 4 | 0 | 100% | |
| _dataframe.py | 55 | 0 | 22 | 0 | 100% | |
| _environment.py | 33 | 1 | 10 | 2 | 96% | 49 |
| _fixes.py | 8 | 0 | 2 | 0 | 100% | |
| _index.py | 15 | 0 | 8 | 0 | 100% | |
| _measure_time.py | 10 | 0 | 0 | 0 | 100% | |
| _parallel.py | 17 | 0 | 0 | 0 | 100% | |
| _patch.py | 21 | 12 | 8 | 8 | 42% | 30, 35–39, 42–43, 46–47, 58, 60 |
| _progress_bar.py | 42 | 4 | 4 | 0 | 90% | 53–54, 64–65 |
| _show_versions.py | 66 | 0 | 22 | 0 | 100% | |
| _skrub.py | 46 | 1 | 8 | 1 | 97% | 37 |
| _testing.py | 128 | 14 | 12 | 2 | 89% | 24, 33, 71–72, 94, 193, 202, 213–218, 220 |
| _uuid.py | 18 | 4 | 10 | 4 | 77% | 7, 17, 23, 26 |
| docscrape.py | 44 | 2 | 12 | 1 | 95% | 32–33 |
| skore/src/skore/_utils/repr | ||||||
| __init__.py | 2 | 0 | 0 | 0 | 100% | |
| base.py | 54 | 0 | 4 | 0 | 100% | |
| data.py | 150 | 0 | 36 | 1 | 100% | |
| html_repr.py | 40 | 0 | 0 | 0 | 100% | |
| markdown.py | 58 | 0 | 22 | 0 | 100% | |
| rich_repr.py | 94 | 0 | 34 | 3 | 100% | |
| utils.py | 20 | 0 | 2 | 0 | 100% | |
| TOTAL | 7777 | 214 | 2142 | 146 | 97% | |
| Tests | Skipped | Failures | Errors | Time |
|---|---|---|---|---|
| 2911 | 3 💤 | 0 ❌ | 0 🔥 | 4m 18s ⏱️ |
There was a problem hiding this comment.
Honestly i have no strong opinion on that.
Making the availability tests in a factory is the same as delegating to each metric.available.
The only benefit is one compute in the factory vs re-doing many computes in each metric.available.
LGTM, wait for the @auguste-probabl review.
auguste-probabl
left a comment
There was a problem hiding this comment.
As I understand it, the point of default_metrics is to allow more flexibility in defining which metrics should be shown: available tells you if the metric is technically computable, whereas default_metrics would tell you what we maintainers recommend showing. The thing is:
- I can't think of a motivating example for this change. What would you like to do with this that you can't do with
available? - If we accept that we need this distinction, I don't see why that information couldn't be defined within each metric class, e.g. with a new
recommendedmethod.
| list of :class:`Metric` | ||
| The metric instances to register, in default display order. | ||
| """ | ||
| ml_task = report._ml_task |
|
The factory with the default choices is different semantic to me. It answers to the following question "From the available metric for the problem, is this particular metric advisable to be used given the problem, data, and model at hand". So in short, |
|
I need to think a bit more for this problem. Basically, what I don't like here is more the language actually: built-in and available are the two terms that I'm concerned with. However, the fact that the metric itself, given the report, can tell you what to do (that could be extended for which reason), is actually nice. I'll think about more about the metric API to see if we can make this distinction between the |
|
@glemaitre if a user adds a metric, do you want |
|
So I'm putting in draft mode because it seems that it is not as urgent as what I thought. |
My last thought where that the If the user is bypassing this, the metric is added in the registry and will be computed. If not bypassing the check, then we raise an error with the reason. In the factory, we could add any available method and make sure it does not raise any error then. |
|
I open #3202 to isolate the part that should not require discussion for the moment. |
Summary
default_metricsfactory with a flatDEFAULT_METRICScatalog.Metric.available→Metric.is_computable(technical: can we compute it?).Metric.discouraged(report) -> str | None(methodological: should we register it?). Defaults skip discouraged metrics silently;metrics.add(..., force=False)raises with the reason unlessforce=True.has_default_scoredisplay filter ontoScore.discouraged, so sklearn's defaultscoreis no longer registered for LogisticRegression / RandomForest / etc.metrics.add(Score(), force=True)to put a discouraged score back intosummarize/get;metrics.score()remains the dedicated escape hatch.Behavior change
Defaults only register metrics that are both computable and not discouraged. User
.addnever checksis_computable; onlydiscouragedcan block it (unlessforce=True).For estimators that use scikit-learn's default
ClassifierMixin.score/RegressorMixin.score,"score"is skipped from the registry for the same reason (usereport.metrics.score(), oradd(Score(), force=True)to put it back).Example
The point of
discouragedis to refuse registering a metric that can be computed but should not be reported for the problem at hand. Override it on your metric class;.addraises with that reason unless the caller passesforce=True:Custom metrics without a
discouragedoverride keep today's behavior:.addregisters them unconditionally (still not filtered byis_computable).Out of scope (flagged, not fixed here)
neg_alias claim insummarize/getdocstrings that lookup does not implementMetric.newsilently dropping some argumentsMetricRow/MetricsSummaryRowduplication'score'listing in accepted design recorddocs/design/0007-show-neg-when-used.md(left as history)Test plan
metrics.score()still work when score is not registeredregistry ↔ reportreference (state size test)main(#3196registry-key row names kept viaresolved.name)