Overview

Dataset statistics

Number of variables4
Number of observations2
Missing cells0
Missing cells (%)0.0%
Duplicate rows0
Duplicate rows (%)0.0%
Total size in memory192.0 B
Average record size in memory96.0 B

Variable types

Categorical4

Alerts

eligible_processes_for_digitalization has constant value "4169" Constant
fiscal_year is highly correlated with completed_digitalized_processesHigh correlation
completed_digitalized_processes is highly correlated with fiscal_yearHigh correlation
fiscal_year is highly correlated with completed_digitalized_processesHigh correlation
completed_digitalized_processes is highly correlated with fiscal_yearHigh correlation
fiscal_year is highly correlated with completed_digitalized_processesHigh correlation
completed_digitalized_processes is highly correlated with fiscal_yearHigh correlation
fiscal_year is highly correlated with completed_digitalized_processes and 2 other fieldsHigh correlation
completed_digitalized_processes is highly correlated with fiscal_year and 2 other fieldsHigh correlation
eligible_processes_for_digitalization is highly correlated with fiscal_year and 2 other fieldsHigh correlation
digitalization_rate is highly correlated with fiscal_year and 2 other fieldsHigh correlation
fiscal_year is uniformly distributed Uniform
completed_digitalized_processes is uniformly distributed Uniform
digitalization_rate is uniformly distributed Uniform
fiscal_year has unique values Unique
completed_digitalized_processes has unique values Unique
digitalization_rate has unique values Unique

Reproduction

Analysis started2026-09-09 04:13:38.711968
Analysis finished2026-09-09 04:13:39.329542
Duration0.62 seconds
Software versionpandas-profiling v3.1.0
Download configurationconfig.json

Variables

fiscal_year
Categorical

HIGH CORRELATION
HIGH CORRELATION
HIGH CORRELATION
HIGH CORRELATION
UNIFORM
UNIQUE

Distinct2
Distinct (%)100.0%
Missing0
Missing (%)0.0%
Memory size144.0 B
2567
1 
2568
1 

Length

Max length4
Median length4
Mean length4
Min length4

Characters and Unicode

Total characters0
Distinct characters0
Distinct categories0 ?
Distinct scripts0 ?
Distinct blocks0 ?
The Unicode Standard assigns character properties to each code point, which can be used to analyse textual variables.

Unique

Unique2 ?
Unique (%)100.0%

Sample

1st row2567
2nd row2568

Common Values

ValueCountFrequency (%)
25671
50.0%
25681
50.0%

Length

2026-09-09T11:13:39.398730image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
Histogram of lengths of the category

Pie chart

2026-09-09T11:13:39.492544image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
ValueCountFrequency (%)
25681
50.0%
25671
50.0%

Most occurring characters

ValueCountFrequency (%)
No values found.

Most occurring categories

ValueCountFrequency (%)
No values found.

Most frequent character per category

Most occurring scripts

ValueCountFrequency (%)
No values found.

Most frequent character per script

Most occurring blocks

ValueCountFrequency (%)
No values found.

Most frequent character per block

completed_digitalized_processes
Categorical

HIGH CORRELATION
HIGH CORRELATION
HIGH CORRELATION
HIGH CORRELATION
UNIFORM
UNIQUE

Distinct2
Distinct (%)100.0%
Missing0
Missing (%)0.0%
Memory size144.0 B
2701
1 
3491
1 

Length

Max length4
Median length4
Mean length4
Min length4

Characters and Unicode

Total characters0
Distinct characters0
Distinct categories0 ?
Distinct scripts0 ?
Distinct blocks0 ?
The Unicode Standard assigns character properties to each code point, which can be used to analyse textual variables.

Unique

Unique2 ?
Unique (%)100.0%

Sample

1st row2701
2nd row3491

Common Values

ValueCountFrequency (%)
27011
50.0%
34911
50.0%

Length

2026-09-09T11:13:39.590223image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
Histogram of lengths of the category

Pie chart

2026-09-09T11:13:39.683422image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
ValueCountFrequency (%)
34911
50.0%
27011
50.0%

Most occurring characters

ValueCountFrequency (%)
No values found.

Most occurring categories

ValueCountFrequency (%)
No values found.

Most frequent character per category

Most occurring scripts

ValueCountFrequency (%)
No values found.

Most frequent character per script

Most occurring blocks

ValueCountFrequency (%)
No values found.

Most frequent character per block

eligible_processes_for_digitalization
Categorical

CONSTANT
HIGH CORRELATION
REJECTED

Distinct1
Distinct (%)50.0%
Missing0
Missing (%)0.0%
Memory size144.0 B
4169
2 

Length

Max length4
Median length4
Mean length4
Min length4

Characters and Unicode

Total characters0
Distinct characters0
Distinct categories0 ?
Distinct scripts0 ?
Distinct blocks0 ?
The Unicode Standard assigns character properties to each code point, which can be used to analyse textual variables.

Unique

Unique0 ?
Unique (%)0.0%

Sample

1st row4169
2nd row4169

Common Values

ValueCountFrequency (%)
41692
100.0%

Length

2026-09-09T11:13:39.779816image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
Histogram of lengths of the category

Pie chart

2026-09-09T11:13:39.870939image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
ValueCountFrequency (%)
41692
100.0%

Most occurring characters

ValueCountFrequency (%)
No values found.

Most occurring categories

ValueCountFrequency (%)
No values found.

Most frequent character per category

Most occurring scripts

ValueCountFrequency (%)
No values found.

Most frequent character per script

Most occurring blocks

ValueCountFrequency (%)
No values found.

Most frequent character per block

digitalization_rate
Categorical

HIGH CORRELATION
UNIFORM
UNIQUE

Distinct2
Distinct (%)100.0%
Missing0
Missing (%)0.0%
Memory size144.0 B
83.74%
1 
64.79%
1 

Length

Max length6
Median length6
Mean length6
Min length6

Characters and Unicode

Total characters0
Distinct characters0
Distinct categories0 ?
Distinct scripts0 ?
Distinct blocks0 ?
The Unicode Standard assigns character properties to each code point, which can be used to analyse textual variables.

Unique

Unique2 ?
Unique (%)100.0%

Sample

1st row64.79%
2nd row83.74%

Common Values

ValueCountFrequency (%)
83.74%1
50.0%
64.79%1
50.0%

Length

2026-09-09T11:13:39.958398image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
Histogram of lengths of the category

Pie chart

2026-09-09T11:13:40.052087image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
ValueCountFrequency (%)
64.791
50.0%
83.741
50.0%

Most occurring characters

ValueCountFrequency (%)
No values found.

Most occurring categories

ValueCountFrequency (%)
No values found.

Most frequent character per category

Most occurring scripts

ValueCountFrequency (%)
No values found.

Most frequent character per script

Most occurring blocks

ValueCountFrequency (%)
No values found.

Most frequent character per block

Correlations

2026-09-09T11:13:40.127403image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/

Spearman's ρ

The Spearman's rank correlation coefficient (ρ) is a measure of monotonic correlation between two variables, and is therefore better in catching nonlinear monotonic correlations than Pearson's r. It's value lies between -1 and +1, -1 indicating total negative monotonic correlation, 0 indicating no monotonic correlation and 1 indicating total positive monotonic correlation.

To calculate ρ for two variables X and Y, one divides the covariance of the rank variables of X and Y by the product of their standard deviations.
2026-09-09T11:13:40.317393image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/

Pearson's r

The Pearson's correlation coefficient (r) is a measure of linear correlation between two variables. It's value lies between -1 and +1, -1 indicating total negative linear correlation, 0 indicating no linear correlation and 1 indicating total positive linear correlation. Furthermore, r is invariant under separate changes in location and scale of the two variables, implying that for a linear function the angle to the x-axis does not affect r.

To calculate r for two variables X and Y, one divides the covariance of X and Y by the product of their standard deviations.
2026-09-09T11:13:40.506202image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/

Kendall's τ

Similarly to Spearman's rank correlation coefficient, the Kendall rank correlation coefficient (τ) measures ordinal association between two variables. It's value lies between -1 and +1, -1 indicating total negative correlation, 0 indicating no correlation and 1 indicating total positive correlation.

To calculate τ for two variables X and Y, one determines the number of concordant and discordant pairs of observations. τ is given by the number of concordant pairs minus the discordant pairs divided by the total number of pairs.
2026-09-09T11:13:40.706357image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/

Cramér's V (φc)

Cramér's V is an association measure for nominal random variables. The coefficient ranges from 0 to 1, with 0 indicating independence and 1 indicating perfect association. The empirical estimators used for Cramér's V have been proved to be biased, even for large samples. We use a bias-corrected measure that has been proposed by Bergsma in 2013 that can be found here.
2026-09-09T11:13:40.883303image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/

Phik (φk)

Phik (φk) is a new and practical correlation coefficient that works consistently between categorical, ordinal and interval variables, captures non-linear dependency and reverts to the Pearson correlation coefficient in case of a bivariate normal input distribution. There is extensive documentation available here.

Missing values

2026-09-09T11:13:39.033623image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
A simple visualization of nullity by column.
2026-09-09T11:13:39.257243image/svg+xmlMatplotlib v3.3.4, https://matplotlib.org/
Nullity matrix is a data-dense display which lets you quickly visually pick out patterns in data completion.

Sample

First rows

fiscal_yearcompleted_digitalized_processeseligible_processes_for_digitalizationdigitalization_rate
025672701416964.79%
125683491416983.74%

Last rows

fiscal_yearcompleted_digitalized_processeseligible_processes_for_digitalizationdigitalization_rate
025672701416964.79%
125683491416983.74%