Thanks to visit codestin.com
Credit goes to github.com

Skip to content

🔒 🤖 CI Update lock files for scipy-dev CI build(s) 🔒 🤖 - #34894

Closed
scikit-learn-bot wants to merge 1 commit into
scikit-learn:mainfrom
scikit-learn-bot:auto-update-lock-files-scipy-dev
Closed

🔒 🤖 CI Update lock files for scipy-dev CI build(s) 🔒 🤖#34894
scikit-learn-bot wants to merge 1 commit into
scikit-learn:mainfrom
scikit-learn-bot:auto-update-lock-files-scipy-dev

Conversation

@scikit-learn-bot

Copy link
Copy Markdown
Contributor

Update lock files.

The update was generated by running:

python build_tools/update_environments_and_lock_files.py --select-tag scipy-dev

Note

If the CI tasks fail, create a new branch based on this PR and add the required fixes to that branch.
Prefer re-running the same command above (with the matching --select-tag) rather than updating all lock files.
Remember to include [scipy-dev] in your commit message so that the optional CI build runs.

@scikit-learn-bot
scikit-learn-bot force-pushed the auto-update-lock-files-scipy-dev branch from 9304b32 to b64dd06 Compare September 7, 2026 05:06
@lesteve

lesteve commented Sep 7, 2026

Copy link
Copy Markdown
Member

September 7: looks like a pandas-dev change that manifests itself in fetch_openml? build log

FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_true[True-pandas-2-dataset_params2-11-38-1] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_true[True-pandas-2-dataset_params3-11-38-1] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_true[True-pandas-40589-dataset_params6-13-72-6] - TypeError: expected string or bytes-like object, got 'bool'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_true[True-pandas-40945-dataset_params11-1309-13-1] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_true[False-pandas-2-dataset_params2-11-38-1] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_true[False-pandas-2-dataset_params3-11-38-1] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_true[False-pandas-40589-dataset_params6-13-72-6] - TypeError: expected string or bytes-like object, got 'bool'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_true[False-pandas-40945-dataset_params11-1309-13-1] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_false[pandas-2-dataset_params2-11-38-1] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_false[pandas-2-dataset_params3-11-38-1] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_as_frame_false[pandas-40589-dataset_params6-13-72-6] - TypeError: expected string or bytes-like object, got 'bool'
FAILED datasets/tests/test_openml.py::test_fetch_openml_consistency_parser[40945] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_equivalence_frame_return_X_y[pandas-2] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_equivalence_frame_return_X_y[pandas-40589] - TypeError: expected string or bytes-like object, got 'bool'
FAILED datasets/tests/test_openml.py::test_fetch_openml_equivalence_array_return_X_y[pandas-40589] - TypeError: expected string or bytes-like object, got 'bool'
FAILED datasets/tests/test_openml.py::test_fetch_openml_types_inference[True-2-pandas-33-2-4] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_types_inference[True-40589-pandas-6-69-3] - TypeError: expected string or bytes-like object, got 'bool'
FAILED datasets/tests/test_openml.py::test_fetch_openml_types_inference[True-40945-pandas-3-3-3] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_types_inference[False-2-pandas-33-2-4] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_types_inference[False-40589-pandas-6-69-3] - TypeError: expected string or bytes-like object, got 'bool'
FAILED datasets/tests/test_openml.py::test_fetch_openml_types_inference[False-40945-pandas-3-3-3] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_verify_checksum[True-pandas-True-False] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_verify_checksum[True-pandas-True-True] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_verify_checksum[False-pandas-True-False] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_verify_checksum[False-pandas-True-True] - TypeError: expected string or bytes-like object, got 'int'
FAILED datasets/tests/test_openml.py::test_fetch_openml_with_ignored_feature[pandas-True] - TypeError: expected string or bytes-like object, got 'bool'
FAILED datasets/tests/test_openml.py::test_fetch_openml_with_ignored_feature[pandas-False] - TypeError: expected string or bytes-like object, got 'bool'
= 27 failed, 39459 passed, 13524 skipped, 227 xfailed, 118 xpassed, 5985 warnings in 407.21s (0:06:47) =
One stack-trace
_____________ test_fetch_openml_with_ignored_feature[pandas-True] ______________
[gw1] linux -- Python 3.14.7 /home/runner/miniconda3/envs/testvenv/bin/python

args = ('https://www.openml.org/data/v1/download/52352/zoo.arff', None)
kw = {'parser': 'pandas', 'output_type': 'numpy', 'openml_columns_info': {'animal': {'index': '0', 'name': 'animal', 'data_... 'true'], ...}, ...}, 'feature_names_to_select': ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...], ...}

    @wraps(f)
    def wrapper(*args, **kw):
        try:
>           return f(*args, **kw)
                   ^^^^^^^^^^^^^^

args       = ('https://www.openml.org/data/v1/download/52352/zoo.arff', None)
data_home  = None
f          = <function _load_arff_response at 0x7f327d8f2a30>
kw         = {'parser': 'pandas', 'output_type': 'numpy', 'openml_columns_info': {'animal': {'index': '0', 'name': 'animal', 'data_... 'true'], ...}, ...}, 'feature_names_to_select': ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...], ...}
no_retry_exception = <class 'pandas.errors.ParserError'>
openml_path = 'data/v1/download/52352/zoo.arff'

../sklearn/datasets/_openml.py:80: 
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ 
../sklearn/datasets/_openml.py:580: in _load_arff_response
    X, y, frame, categories = _open_url_and_load_gzip_file(
        ParserError = <class 'pandas.errors.ParserError'>
        _open_url_and_load_gzip_file = <function _load_arff_response.<locals>._open_url_and_load_gzip_file at 0x7f32589dee50>
        actual_md5_checksum = 'a84cd1e10cc766cd32fd329c2ada36ac'
        arff_params = {'parser': 'pandas', 'output_type': 'numpy', 'openml_columns_info': {'animal': {'index': '0', 'name': 'animal', 'data_... 'true'], ...}, ...}, 'feature_names_to_select': ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...], ...}
        chunk      = b'ammal\nraccoon,true,false,false,true,false,false,true,true,true,true,false,false,4,true,false,true,mammal\nreindeer,...,false,invertebrate\nwren,false,true,true,false,true,false,false,false,true,true,false,false,2,true,false,false,bird\n'
        data_home  = None
        delay      = 1.0
        feature_names_to_select = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        gzip_file  = <gzip on 0x7f3258b15d20>
        md5        = <md5 _hashlib.HASH object @ 0x7f325b14c750>
        md5_checksum = 'a84cd1e10cc766cd32fd329c2ada36ac'
        n_retries  = 3
        openml_columns_info = {'animal': {'index': '0', 'name': 'animal', 'data_type': 'nominal', 'nominal_value': ['aardvark', 'antelope', 'bass', ...'], ...}, 'eggs': {'index': '3', 'name': 'eggs', 'data_type': 'nominal', 'nominal_value': ['false', 'true'], ...}, ...}
        output_type = 'numpy'
        parser     = 'pandas'
        read_csv_kwargs = None
        shape      = (101, 18)
        target_names_to_select = ['type']
        url        = 'https://www.openml.org/data/v1/download/52352/zoo.arff'
../sklearn/datasets/_openml.py:568: in _open_url_and_load_gzip_file
    return load_arff_from_gzip_file(gzip_file, **arff_params)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        arff_params = {'parser': 'pandas', 'output_type': 'numpy', 'openml_columns_info': {'animal': {'index': '0', 'name': 'animal', 'data_... 'true'], ...}, ...}, 'feature_names_to_select': ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...], ...}
        data_home  = None
        delay      = 1.0
        gzip_file  = <gzip on 0x7f3258b15510>
        n_retries  = 3
        url        = 'https://www.openml.org/data/v1/download/52352/zoo.arff'
../sklearn/datasets/_arff_parser.py:533: in load_arff_from_gzip_file
    return _pandas_arff_parser(
        feature_names_to_select = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        gzip_file  = <gzip on 0x7f3258b15510>
        openml_columns_info = {'animal': {'index': '0', 'name': 'animal', 'data_type': 'nominal', 'nominal_value': ['aardvark', 'antelope', 'bass', ...'], ...}, 'eggs': {'index': '3', 'name': 'eggs', 'data_type': 'nominal', 'nominal_value': ['false', 'true'], ...}, ...}
        output_type = 'numpy'
        parser     = 'pandas'
        read_csv_kwargs = {}
        shape      = (101, 18)
        target_names_to_select = ['type']
../sklearn/datasets/_arff_parser.py:447: in _pandas_arff_parser
    frame[col] = frame[col].cat.rename_categories(strip_single_quotes)
                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        categorical_columns = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        col        = 'hair'
        column_dtype = 'nominal'
        columns_to_keep = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        columns_to_select = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        default_read_csv_kwargs = {'header': None, 'index_col': False, 'na_values': ['?'], 'keep_default_na': False, ...}
        dtypes     = {'animal': 'category', 'hair': 'category', 'feathers': 'category', 'eggs': 'category', ...}
        dtypes_positional = {0: 'category', 1: 'category', 2: 'category', 3: 'category', ...}
        feature_names_to_select = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        frame      =       hair feathers   eggs   milk  ...   tail domestic catsize          type
0     True    False  False   True  ...  F...lse  invertebrate
100  False     True   True  False  ...   True    False   False          bird

[101 rows x 17 columns]
        gzip_file  = <gzip on 0x7f3258b15510>
        line       = b'@data\n'
        name       = 'type'
        openml_columns_info = {'animal': {'index': '0', 'name': 'animal', 'data_type': 'nominal', 'nominal_value': ['aardvark', 'antelope', 'bass', ...'], ...}, 'eggs': {'index': '3', 'name': 'eggs', 'data_type': 'nominal', 'nominal_value': ['false', 'true'], ...}, ...}
        output_arrays_type = 'numpy'
        pd         = <module 'pandas' from '/home/runner/miniconda3/envs/testvenv/lib/python3.14/site-packages/pandas/__init__.py'>
        read_csv_kwargs = {'header': None, 'index_col': False, 'na_values': ['?'], 'keep_default_na': False, ...}
        single_quote_pattern = re.compile("^'(?P<contents>.*)'$")
        strip_single_quotes = <function _pandas_arff_parser.<locals>.strip_single_quotes at 0x7f325855c040>
        target_names_to_select = ['type']
../../../../miniconda3/envs/testvenv/lib/python3.14/site-packages/pandas/core/accessor.py:128: in f
    return self._delegate_method(name, *args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        args       = (<function _pandas_arff_parser.<locals>.strip_single_quotes at 0x7f325855c040>,)
        kwargs     = {}
        name       = 'rename_categories'
        self       = <pandas.core.arrays.categorical.CategoricalAccessor object at 0x7f3258b4cd70>
../../../../miniconda3/envs/testvenv/lib/python3.14/site-packages/pandas/core/arrays/categorical.py:3252: in _delegate_method
    res = method(*args, **kwargs)
          ^^^^^^^^^^^^^^^^^^^^^^^
        Series     = <class 'pandas.Series'>
        args       = (<function _pandas_arff_parser.<locals>.strip_single_quotes at 0x7f325855c040>,)
        kwargs     = {}
        method     = <bound method Categorical.rename_categories of [True, True, False, True, True, ..., True, True, True, False, False]
Length: 101
Categories (2, bool): [False, True]>
        name       = 'rename_categories'
        self       = <pandas.core.arrays.categorical.CategoricalAccessor object at 0x7f3258b4cd70>
../../../../miniconda3/envs/testvenv/lib/python3.14/site-packages/pandas/core/arrays/categorical.py:1429: in rename_categories
    new_categories = [new_categories(item) for item in self.categories]
                      ^^^^^^^^^^^^^^^^^^^^
        new_categories = <function _pandas_arff_parser.<locals>.strip_single_quotes at 0x7f325855c040>
        self       = [True, True, False, True, True, ..., True, True, True, False, False]
Length: 101
Categories (2, bool): [False, True]
../sklearn/datasets/_arff_parser.py:435: in strip_single_quotes
    match = re.search(single_quote_pattern, input_string)
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        input_string = False
        single_quote_pattern = re.compile("^'(?P<contents>.*)'$")
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ 

pattern = re.compile("^'(?P<contents>.*)'$"), string = False, flags = 0

    def search(pattern, string, flags=0):
        """Scan through string looking for a match to the pattern, returning
        a Match object, or None if no match was found."""
>       return _compile(pattern, flags).search(string)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
E       TypeError: expected string or bytes-like object, got 'bool'

flags      = 0
pattern    = re.compile("^'(?P<contents>.*)'$")
string     = False

../../../../miniconda3/envs/testvenv/lib/python3.14/re/__init__.py:177: TypeError

During handling of the above exception, another exception occurred:

monkeypatch = <_pytest.monkeypatch.MonkeyPatch object at 0x7f3269dbb460>
gzip_response = True, parser = 'pandas'

    @pytest.mark.parametrize("gzip_response", [True, False])
    @pytest.mark.parametrize("parser", ("liac-arff", "pandas"))
    def test_fetch_openml_with_ignored_feature(monkeypatch, gzip_response, parser):
        """Check that we can load the "zoo" dataset.
        Non-regression test for:
        https://github.com/scikit-learn/scikit-learn/issues/14340
        """
        if parser == "pandas":
            pytest.importorskip("pandas")
        data_id = 62
        _monkey_patch_webbased_functions(monkeypatch, data_id, gzip_response)
    
>       dataset = sklearn.datasets.fetch_openml(
            data_id=data_id, cache=False, as_frame=False, parser=parser
        )

data_id    = 62
gzip_response = True
monkeypatch = <_pytest.monkeypatch.MonkeyPatch object at 0x7f3269dbb460>
parser     = 'pandas'

../sklearn/datasets/tests/test_openml.py:1617: 
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ 
../sklearn/utils/_param_validation.py:218: in wrapper
    return func(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^
        args       = ()
        func       = <function fetch_openml at 0x7f327d8f2f00>
        func_sig   = <Signature (name: str | None = None, *, version: str | int = 'active', data_id: int | None = None, data_home: str | os...tr | bool = 'auto', n_retries: int = 3, delay: float = 1.0, parser: str = 'auto', read_csv_kwargs: Dict | None = None)>
        global_skip_validation = False
        kwargs     = {'data_id': 62, 'cache': False, 'as_frame': False, 'parser': 'pandas'}
        parameter_constraints = {'name': [<class 'str'>, None], 'version': [<sklearn.utils._param_validation.Interval object at 0x7f327d84b460>, <skle...m_validation.Interval object at 0x7f327d84b540>, None], 'data_home': [<class 'str'>, <class 'os.PathLike'>, None], ...}
        params     = {'name': None, 'version': 'active', 'data_id': 62, 'data_home': None, ...}
        prefer_skip_nested_validation = True
        to_ignore  = ['self', 'cls']
../sklearn/datasets/_openml.py:1161: in fetch_openml
    bunch = _download_data_to_bunch(
        as_frame   = False
        cache      = False
        data_columns = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        data_description = {'id': '62', 'name': 'zoo', 'version': '1', 'description': '**Author**: Richard S. Forsyth   \n**Source**: [UCI](https...nd one of "girl"!\n* feature \'animal\' is an identifier (though not unique) and should be ignored when modeling', ...}
        data_home  = None
        data_id    = 62
        data_qualities = [{'name': 'AutoCorrelation', 'value': '0.35'}, {'name': 'ClassEntropy', 'value': '2.3905596822940396'}, {'name': 'Dime... {'name': 'MajorityClassPercentage', 'value': '40.5940594059406'}, {'name': 'MajorityClassSize', 'value': '41.0'}, ...]
        delay      = 1.0
        feature    = {'index': '17', 'name': 'type', 'data_type': 'nominal', 'nominal_value': ['mammal', 'bird', 'reptile', 'fish', 'amphibian', 'insect', ...], ...}
        features_list = [{'index': '0', 'name': 'animal', 'data_type': 'nominal', 'nominal_value': ['aardvark', 'antelope', 'bass', 'bear', 'b...true'], ...}, {'index': '5', 'name': 'airborne', 'data_type': 'nominal', 'nominal_value': ['false', 'true'], ...}, ...]
        n_retries  = 3
        name       = None
        parser     = 'pandas'
        parser_    = 'pandas'
        read_csv_kwargs = None
        return_X_y = False
        return_sparse = False
        shape      = (101, 18)
        target_column = 'default-target'
        target_columns = ['type']
        url        = 'https://www.openml.org/data/v1/download/52352/zoo.arff'
        version    = 'active'
../sklearn/datasets/_openml.py:715: in _download_data_to_bunch
    X, y, frame, categories = _retry_with_clean_cache(
        ParserError = <class 'pandas.errors.ParserError'>
        as_frame   = False
        column_info = {'index': '17', 'name': 'type', 'data_type': 'nominal', 'nominal_value': ['mammal', 'bird', 'reptile', 'fish', 'amphibian', 'insect', ...], ...}
        data_columns = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        data_home  = None
        delay      = 1.0
        features_dict = {'animal': {'index': '0', 'name': 'animal', 'data_type': 'nominal', 'nominal_value': ['aardvark', 'antelope', 'bass', ...'], ...}, 'eggs': {'index': '3', 'name': 'eggs', 'data_type': 'nominal', 'nominal_value': ['false', 'true'], ...}, ...}
        md5_checksum = 'a84cd1e10cc766cd32fd329c2ada36ac'
        n_missing_values = 0
        n_retries  = 3
        name       = 'type'
        no_retry_exception = <class 'pandas.errors.ParserError'>
        openml_columns_info = [{'index': '0', 'name': 'animal', 'data_type': 'nominal', 'nominal_value': ['aardvark', 'antelope', 'bass', 'bear', 'b...true'], ...}, {'index': '5', 'name': 'airborne', 'data_type': 'nominal', 'nominal_value': ['false', 'true'], ...}, ...]
        output_type = 'numpy'
        parser     = 'pandas'
        read_csv_kwargs = None
        shape      = (101, 18)
        sparse     = False
        target_columns = ['type']
        url        = 'https://www.openml.org/data/v1/download/52352/zoo.arff'
../sklearn/datasets/_openml.py:100: in wrapper
    return f(*args, **kw)
           ^^^^^^^^^^^^^^
        args       = ('https://www.openml.org/data/v1/download/52352/zoo.arff', None)
        data_home  = None
        f          = <function _load_arff_response at 0x7f327d8f2a30>
        kw         = {'parser': 'pandas', 'output_type': 'numpy', 'openml_columns_info': {'animal': {'index': '0', 'name': 'animal', 'data_... 'true'], ...}, ...}, 'feature_names_to_select': ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...], ...}
        no_retry_exception = <class 'pandas.errors.ParserError'>
        openml_path = 'data/v1/download/52352/zoo.arff'
../sklearn/datasets/_openml.py:580: in _load_arff_response
    X, y, frame, categories = _open_url_and_load_gzip_file(
        ParserError = <class 'pandas.errors.ParserError'>
        _open_url_and_load_gzip_file = <function _load_arff_response.<locals>._open_url_and_load_gzip_file at 0x7f325855c460>
        actual_md5_checksum = 'a84cd1e10cc766cd32fd329c2ada36ac'
        arff_params = {'parser': 'pandas', 'output_type': 'numpy', 'openml_columns_info': {'animal': {'index': '0', 'name': 'animal', 'data_... 'true'], ...}, ...}, 'feature_names_to_select': ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...], ...}
        chunk      = b'ammal\nraccoon,true,false,false,true,false,false,true,true,true,true,false,false,4,true,false,true,mammal\nreindeer,...,false,invertebrate\nwren,false,true,true,false,true,false,false,false,true,true,false,false,2,true,false,false,bird\n'
        data_home  = None
        delay      = 1.0
        feature_names_to_select = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        gzip_file  = <gzip on 0x7f3258b01cf0>
        md5        = <md5 _hashlib.HASH object @ 0x7f325b14e950>
        md5_checksum = 'a84cd1e10cc766cd32fd329c2ada36ac'
        n_retries  = 3
        openml_columns_info = {'animal': {'index': '0', 'name': 'animal', 'data_type': 'nominal', 'nominal_value': ['aardvark', 'antelope', 'bass', ...'], ...}, 'eggs': {'index': '3', 'name': 'eggs', 'data_type': 'nominal', 'nominal_value': ['false', 'true'], ...}, ...}
        output_type = 'numpy'
        parser     = 'pandas'
        read_csv_kwargs = None
        shape      = (101, 18)
        target_names_to_select = ['type']
        url        = 'https://www.openml.org/data/v1/download/52352/zoo.arff'
../sklearn/datasets/_openml.py:568: in _open_url_and_load_gzip_file
    return load_arff_from_gzip_file(gzip_file, **arff_params)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        arff_params = {'parser': 'pandas', 'output_type': 'numpy', 'openml_columns_info': {'animal': {'index': '0', 'name': 'animal', 'data_... 'true'], ...}, ...}, 'feature_names_to_select': ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...], ...}
        data_home  = None
        delay      = 1.0
        gzip_file  = <gzip on 0x7f3258b01c60>
        n_retries  = 3
        url        = 'https://www.openml.org/data/v1/download/52352/zoo.arff'
../sklearn/datasets/_arff_parser.py:533: in load_arff_from_gzip_file
    return _pandas_arff_parser(
        feature_names_to_select = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        gzip_file  = <gzip on 0x7f3258b01c60>
        openml_columns_info = {'animal': {'index': '0', 'name': 'animal', 'data_type': 'nominal', 'nominal_value': ['aardvark', 'antelope', 'bass', ...'], ...}, 'eggs': {'index': '3', 'name': 'eggs', 'data_type': 'nominal', 'nominal_value': ['false', 'true'], ...}, ...}
        output_type = 'numpy'
        parser     = 'pandas'
        read_csv_kwargs = {}
        shape      = (101, 18)
        target_names_to_select = ['type']
../sklearn/datasets/_arff_parser.py:447: in _pandas_arff_parser
    frame[col] = frame[col].cat.rename_categories(strip_single_quotes)
                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        categorical_columns = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        col        = 'hair'
        column_dtype = 'nominal'
        columns_to_keep = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        columns_to_select = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        default_read_csv_kwargs = {'header': None, 'index_col': False, 'na_values': ['?'], 'keep_default_na': False, ...}
        dtypes     = {'animal': 'category', 'hair': 'category', 'feathers': 'category', 'eggs': 'category', ...}
        dtypes_positional = {0: 'category', 1: 'category', 2: 'category', 3: 'category', ...}
        feature_names_to_select = ['hair', 'feathers', 'eggs', 'milk', 'airborne', 'aquatic', ...]
        frame      =       hair feathers   eggs   milk  ...   tail domestic catsize          type
0     True    False  False   True  ...  F...lse  invertebrate
100  False     True   True  False  ...   True    False   False          bird

[101 rows x 17 columns]
        gzip_file  = <gzip on 0x7f3258b01c60>
        line       = b'@data\n'
        name       = 'type'
        openml_columns_info = {'animal': {'index': '0', 'name': 'animal', 'data_type': 'nominal', 'nominal_value': ['aardvark', 'antelope', 'bass', ...'], ...}, 'eggs': {'index': '3', 'name': 'eggs', 'data_type': 'nominal', 'nominal_value': ['false', 'true'], ...}, ...}
        output_arrays_type = 'numpy'
        pd         = <module 'pandas' from '/home/runner/miniconda3/envs/testvenv/lib/python3.14/site-packages/pandas/__init__.py'>
        read_csv_kwargs = {'header': None, 'index_col': False, 'na_values': ['?'], 'keep_default_na': False, ...}
        single_quote_pattern = re.compile("^'(?P<contents>.*)'$")
        strip_single_quotes = <function _pandas_arff_parser.<locals>.strip_single_quotes at 0x7f325855c300>
        target_names_to_select = ['type']
../../../../miniconda3/envs/testvenv/lib/python3.14/site-packages/pandas/core/accessor.py:128: in f
    return self._delegate_method(name, *args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        args       = (<function _pandas_arff_parser.<locals>.strip_single_quotes at 0x7f325855c300>,)
        kwargs     = {}
        name       = 'rename_categories'
        self       = <pandas.core.arrays.categorical.CategoricalAccessor object at 0x7f325b0f4ec0>
../../../../miniconda3/envs/testvenv/lib/python3.14/site-packages/pandas/core/arrays/categorical.py:3252: in _delegate_method
    res = method(*args, **kwargs)
          ^^^^^^^^^^^^^^^^^^^^^^^
        Series     = <class 'pandas.Series'>
        args       = (<function _pandas_arff_parser.<locals>.strip_single_quotes at 0x7f325855c300>,)
        kwargs     = {}
        method     = <bound method Categorical.rename_categories of [True, True, False, True, True, ..., True, True, True, False, False]
Length: 101
Categories (2, bool): [False, True]>
        name       = 'rename_categories'
        self       = <pandas.core.arrays.categorical.CategoricalAccessor object at 0x7f325b0f4ec0>
../../../../miniconda3/envs/testvenv/lib/python3.14/site-packages/pandas/core/arrays/categorical.py:1429: in rename_categories
    new_categories = [new_categories(item) for item in self.categories]
                      ^^^^^^^^^^^^^^^^^^^^
        new_categories = <function _pandas_arff_parser.<locals>.strip_single_quotes at 0x7f325855c300>
        self       = [True, True, False, True, True, ..., True, True, True, False, False]
Length: 101
Categories (2, bool): [False, True]
../sklearn/datasets/_arff_parser.py:435: in strip_single_quotes
    match = re.search(single_quote_pattern, input_string)
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
        input_string = False
        single_quote_pattern = re.compile("^'(?P<contents>.*)'$")
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ 

pattern = re.compile("^'(?P<contents>.*)'$"), string = False, flags = 0

    def search(pattern, string, flags=0):
        """Scan through string looking for a match to the pattern, returning
        a Match object, or None if no match was found."""
>       return _compile(pattern, flags).search(string)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
E       TypeError: expected string or bytes-like object, got 'bool'

flags      = 0
pattern    = re.compile("^'(?P<contents>.*)'$")
string     = False

../../../../miniconda3/envs/testvenv/lib/python3.14/re/__init__.py:177: TypeError

@DeaMariaLeon

Copy link
Copy Markdown
Member

I reproduced the bug locally.
I'll keep working on this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants