-Se utiliza el siguiente formato de fecha %Month/%Day/%Year
Explicación de variables operacionales:
Timestamp : Fecha al día de los datos captados
T_H : Tonelaje que se procesa por día
Posion_post : Altura de carrera en el poste en
[mm]
CSS : Close Side Setting (permite controlar la
granulometría triturado)
Presion_Hidroset : Presion de aceite para elevar poste
en PSI)
Corriente : Corriente del motor (Amperes)
T_aceite_motor : Grados Celsius (°C)
F80_CV_001 : Tamaño de partícula en correa alimentada
por chancador [mm]
F80_CV_701 : Tamaño de partícula en correa alimentada
por chancador [mm]
F80_CV_005 : Tamaño de partícula en correa alimentada
por chancador [mm]
Campaña : Número de campaña
Tonelaje_acumulado_CH1 : Tonelaje acumulado de mineral
procesado por cóncavas (Toneladas)
Explicación de variables extraidas de reportes:
date_Time : Fecha de captura de imagenes 3d
Numcampana_med : Numero de campaña y número de
mediciones en la campaña (X_Y)
diametro_Mto : Unidad en pulgadas
dias_acum_Ccv : dias de avance de campaña según
registros (días)
ton_acum_Ccv : tonelaje acumulado según informacion
proporcionada por cliente (Toneladas)
pos_med_Ccv : Posicion de medición de cóncavas
inferior
remanente_Ccv : Espesor remanente de cóncavas promedio
[mm]
desgaste_Ccv : Desgaste promedio de cóncavas [mm]
tasa_desgaste_Ccv_acum : tasa de desgaste mm/Mton
acumulada
aplan_Ccv : Ángulo de inclinación de cóncavas con
respecto a fondo vertical [°]
pos_aplan_Ccv: posición de la medicion del ángulo de
cóncava con respecto al pie de cóncava [mm].
Librerias necesarias para el analisis
library(ranger)
library(glmnet)
library(purrr)
library(pheatmap)
library(lubridate)
library(dplyr)
library(tidyr)
library(readr)
library(readxl)
library(ggplot2)
library(tidymodels)
library(data.table)
library(tibble)
library(openxlsx)
library(knitr)
library(corrplot)
library(ggpubr)
library(ggrepel)
library(broom)
library(reshape2)
library(caret)
library(e1071)
library(kableExtra)
library(webshot2)
library(Amelia)
library(pander)
library(VIM)
library(mice)
library(car) library(gridExtra)
library(zoo)
library(reticulate) } library(moments)
library(DT)
library(rlang)
library(scales)
library(rpart)
library(rpart.plot)
library(tidyverse)
library(caret)
library(FactoMineR)
library(factoextra)
library(nloptr)
library(Metrics) library(randomForest)
library(sensitivity)
## Cargando paquete requerido: Matrix
## Loaded glmnet 4.1-8
##
## Adjuntando el paquete: 'lubridate'
## The following objects are masked from 'package:base':
##
## date, intersect, setdiff, union
##
## Adjuntando el paquete: 'dplyr'
## The following objects are masked from 'package:stats':
##
## filter, lag
## The following objects are masked from 'package:base':
##
## intersect, setdiff, setequal, union
##
## Adjuntando el paquete: 'tidyr'
## The following objects are masked from 'package:Matrix':
##
## expand, pack, unpack
## ── Attaching packages ────────────────────────────────────── tidymodels 1.2.0 ──
## ✔ broom 1.0.5 ✔ rsample 1.2.1
## ✔ dials 1.2.1 ✔ tibble 3.2.1
## ✔ infer 1.0.7 ✔ tune 1.2.1
## ✔ modeldata 1.3.0 ✔ workflows 1.1.4
## ✔ parsnip 1.2.1 ✔ workflowsets 1.1.0
## ✔ recipes 1.0.10 ✔ yardstick 1.3.1
## ── Conflicts ───────────────────────────────────────── tidymodels_conflicts() ──
## ✖ scales::discard() masks purrr::discard()
## ✖ tidyr::expand() masks Matrix::expand()
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ✖ tidyr::pack() masks Matrix::pack()
## ✖ yardstick::spec() masks readr::spec()
## ✖ recipes::step() masks stats::step()
## ✖ tidyr::unpack() masks Matrix::unpack()
## ✖ recipes::update() masks Matrix::update(), stats::update()
## • Dig deeper into tidy modeling with R at https://www.tmwr.org
##
## Adjuntando el paquete: 'data.table'
## The following objects are masked from 'package:dplyr':
##
## between, first, last
## The following objects are masked from 'package:lubridate':
##
## hour, isoweek, mday, minute, month, quarter, second, wday, week,
## yday, year
## The following object is masked from 'package:purrr':
##
## transpose
## corrplot 0.92 loaded
##
## Adjuntando el paquete: 'reshape2'
## The following objects are masked from 'package:data.table':
##
## dcast, melt
## The following object is masked from 'package:tidyr':
##
## smiths
## Cargando paquete requerido: lattice
##
## Adjuntando el paquete: 'caret'
## The following objects are masked from 'package:yardstick':
##
## precision, recall, sensitivity, specificity
## The following object is masked from 'package:purrr':
##
## lift
##
## Adjuntando el paquete: 'e1071'
## The following object is masked from 'package:tune':
##
## tune
## The following object is masked from 'package:rsample':
##
## permutations
## The following object is masked from 'package:parsnip':
##
## tune
##
## Adjuntando el paquete: 'kableExtra'
## The following object is masked from 'package:dplyr':
##
## group_rows
## Cargando paquete requerido: Rcpp
##
## Adjuntando el paquete: 'Rcpp'
## The following object is masked from 'package:rsample':
##
## populate
## ##
## ## Amelia II: Multiple Imputation
## ## (Version 1.8.2, built: 2024-04-10)
## ## Copyright (C) 2005-2024 James Honaker, Gary King and Matthew Blackwell
## ## Refer to http://gking.harvard.edu/amelia/ for more information
## ##
## Cargando paquete requerido: colorspace
## Cargando paquete requerido: grid
## VIM is ready to use.
## Suggestions and bug-reports can be submitted at: https://github.com/statistikat/VIM/issues
##
## Adjuntando el paquete: 'VIM'
## The following object is masked from 'package:recipes':
##
## prepare
## The following object is masked from 'package:datasets':
##
## sleep
##
## Adjuntando el paquete: 'mice'
## The following object is masked from 'package:stats':
##
## filter
## The following objects are masked from 'package:base':
##
## cbind, rbind
## Cargando paquete requerido: carData
##
## Adjuntando el paquete: 'car'
## The following object is masked from 'package:dplyr':
##
## recode
## The following object is masked from 'package:purrr':
##
## some
##
## Adjuntando el paquete: 'gridExtra'
## The following object is masked from 'package:dplyr':
##
## combine
##
## Adjuntando el paquete: 'zoo'
## The following objects are masked from 'package:data.table':
##
## yearmon, yearqtr
## The following objects are masked from 'package:base':
##
## as.Date, as.Date.numeric
##
## Adjuntando el paquete: 'moments'
## The following objects are masked from 'package:e1071':
##
## kurtosis, moment, skewness
##
## Adjuntando el paquete: 'rlang'
## The following object is masked from 'package:data.table':
##
## :=
## The following objects are masked from 'package:purrr':
##
## %@%, flatten, flatten_chr, flatten_dbl, flatten_int, flatten_lgl,
## flatten_raw, invoke, splice
##
## Adjuntando el paquete: 'rpart'
## The following object is masked from 'package:dials':
##
## prune
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ forcats 1.0.0 ✔ stringr 1.5.1
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ rlang::%@%() masks purrr::%@%()
## ✖ rlang:::=() masks data.table:::=()
## ✖ data.table::between() masks dplyr::between()
## ✖ scales::col_factor() masks readr::col_factor()
## ✖ gridExtra::combine() masks dplyr::combine()
## ✖ scales::discard() masks purrr::discard()
## ✖ tidyr::expand() masks Matrix::expand()
## ✖ mice::filter() masks dplyr::filter(), stats::filter()
## ✖ data.table::first() masks dplyr::first()
## ✖ stringr::fixed() masks recipes::fixed()
## ✖ rlang::flatten() masks purrr::flatten()
## ✖ rlang::flatten_chr() masks purrr::flatten_chr()
## ✖ rlang::flatten_dbl() masks purrr::flatten_dbl()
## ✖ rlang::flatten_int() masks purrr::flatten_int()
## ✖ rlang::flatten_lgl() masks purrr::flatten_lgl()
## ✖ rlang::flatten_raw() masks purrr::flatten_raw()
## ✖ kableExtra::group_rows() masks dplyr::group_rows()
## ✖ data.table::hour() masks lubridate::hour()
## ✖ rlang::invoke() masks purrr::invoke()
## ✖ data.table::isoweek() masks lubridate::isoweek()
## ✖ dplyr::lag() masks stats::lag()
## ✖ data.table::last() masks dplyr::last()
## ✖ caret::lift() masks purrr::lift()
## ✖ data.table::mday() masks lubridate::mday()
## ✖ data.table::minute() masks lubridate::minute()
## ✖ data.table::month() masks lubridate::month()
## ✖ tidyr::pack() masks Matrix::pack()
## ✖ data.table::quarter() masks lubridate::quarter()
## ✖ car::recode() masks dplyr::recode()
## ✖ data.table::second() masks lubridate::second()
## ✖ car::some() masks purrr::some()
## ✖ yardstick::spec() masks readr::spec()
## ✖ rlang::splice() masks purrr::splice()
## ✖ data.table::transpose() masks purrr::transpose()
## ✖ tidyr::unpack() masks Matrix::unpack()
## ✖ data.table::wday() masks lubridate::wday()
## ✖ data.table::week() masks lubridate::week()
## ✖ data.table::yday() masks lubridate::yday()
## ✖ data.table::year() masks lubridate::year()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
## Welcome! Want to learn more? See two factoextra-related books at https://goo.gl/ve3WBa
##
##
## Adjuntando el paquete: 'Metrics'
##
##
## The following object is masked from 'package:rlang':
##
## ll
##
##
## The following objects are masked from 'package:caret':
##
## precision, recall
##
##
## The following objects are masked from 'package:yardstick':
##
## accuracy, mae, mape, mase, precision, recall, rmse, smape
##
##
## randomForest 4.7-1.1
##
## Type rfNews() to see new features/changes/bug fixes.
##
##
## Adjuntando el paquete: 'randomForest'
##
##
## The following object is masked from 'package:gridExtra':
##
## combine
##
##
## The following object is masked from 'package:ggplot2':
##
## margin
##
##
## The following object is masked from 'package:dplyr':
##
## combine
##
##
## The following object is masked from 'package:ranger':
##
## importance
##
##
## Registered S3 method overwritten by 'sensitivity':
## method from
## print.src dplyr
##
##
## Adjuntando el paquete: 'sensitivity'
##
##
## The following object is masked from 'package:car':
##
## scatterplot
##
##
## The following object is masked from 'package:tidyr':
##
## extract
##
##
## The following object is masked from 'package:dplyr':
##
## src
# Cargar directorio de trabajo y setear semilla
set.seed(19831118)
getwd()
## [1] "C:/Users/vdiaz/Documents/TESIS"
setwd("C:/Users/vdiaz/Documents/TESIS")
#Carga de archivo de datos
datosOrig <- read_excel("Datos.xlsx")
head(datosOrig, 30)
## # A tibble: 30 × 23
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset
## <dttm> <dbl> <dbl> <dbl> <dbl>
## 1 2019-01-01 00:00:00 154979. 202. 7.5 258.
## 2 2019-01-02 00:00:00 146933. 204. 7.5 251.
## 3 2019-01-03 00:00:00 127159. 193. 7.5 244.
## 4 2019-01-04 00:00:00 113665. 218. 7.5 248.
## 5 2019-01-05 00:00:00 115160. 228. 7.5 241.
## 6 2019-01-06 00:00:00 125231. 235. 7.5 262.
## 7 2019-01-07 00:00:00 115360. 233. 7.54 285.
## 8 2019-01-08 00:00:00 94863. 228. 8 268.
## 9 2019-01-09 00:00:00 87410. 232. 8 266.
## 10 2019-01-10 00:00:00 95211. 244. 8 276.
## # ℹ 20 more rows
## # ℹ 18 more variables: Corriente <dbl>, T_aceite_motor <dbl>, F80_CV_001 <dbl>,
## # F80_CV_701 <dbl>, F80_CV_005 <dbl>, Campaña <chr>,
## # Tonelaje_acumulado_CH1 <dbl>, date_Time <dttm>, Numcampana_med <chr>,
## # diametro_Mto <chr>, dias_acum_Ccv <dbl>, ton_acum_Ccv <dbl>,
## # pos_med_Ccv <dbl>, remanente_Ccv <chr>, desgaste_Ccv <chr>,
## # tasa_desgaste_Ccv_acum <chr>, aplan_Ccv <dbl>, pos_aplan_Ccv <dbl>
str(datosOrig)
## tibble [1,145 × 23] (S3: tbl_df/tbl/data.frame)
## $ Timestamp : POSIXct[1:1145], format: "2019-01-01" "2019-01-02" ...
## $ T_H : num [1:1145] 154979 146933 127159 113665 115160 ...
## $ Posion_poste_manto : num [1:1145] 202 204 193 218 228 ...
## $ CSS : num [1:1145] 7.5 7.5 7.5 7.5 7.5 ...
## $ Presion_Hidroset : num [1:1145] 258 251 244 248 241 ...
## $ Corriente : num [1:1145] 103 102 101 103 102 ...
## $ T_aceite_motor : num [1:1145] 37.7 36.8 38.8 54.2 51.3 ...
## $ F80_CV_001 : num [1:1145] 2.73 2.58 2.51 2.58 2.08 ...
## $ F80_CV_701 : num [1:1145] 0.01 0 0.315 2.459 2.214 ...
## $ F80_CV_005 : num [1:1145] 2.1 1.6 1.85 2.89 2.97 ...
## $ Campaña : chr [1:1145] "C4" "C4" "C4" "C4" ...
## $ Tonelaje_acumulado_CH1: num [1:1145] 0.155 0.302 0.429 0.543 0.658 ...
## $ date_Time : POSIXct[1:1145], format: NA NA ...
## $ Numcampana_med : chr [1:1145] NA NA NA NA ...
## $ diametro_Mto : chr [1:1145] NA NA NA NA ...
## $ dias_acum_Ccv : num [1:1145] NA NA NA NA NA NA NA NA NA NA ...
## $ ton_acum_Ccv : num [1:1145] NA NA NA NA NA NA NA NA NA NA ...
## $ pos_med_Ccv : num [1:1145] NA NA NA NA NA NA NA NA NA NA ...
## $ remanente_Ccv : chr [1:1145] NA NA NA NA ...
## $ desgaste_Ccv : chr [1:1145] NA NA NA NA ...
## $ tasa_desgaste_Ccv_acum: chr [1:1145] NA NA NA NA ...
## $ aplan_Ccv : num [1:1145] NA NA NA NA NA NA NA NA NA NA ...
## $ pos_aplan_Ccv : num [1:1145] NA NA NA NA NA NA NA NA NA NA ...
datos1 <- data.frame(datosOrig)
head(datos1)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset Corriente
## 1 2019-01-01 154979.3 202.2061 7.5 258.3561 102.9613
## 2 2019-01-02 146932.7 204.2289 7.5 251.4526 102.3415
## 3 2019-01-03 127159.3 193.3790 7.5 244.4037 101.0957
## 4 2019-01-04 113665.4 217.8944 7.5 248.4902 103.3990
## 5 2019-01-05 115159.8 227.8408 7.5 241.4058 101.6643
## 6 2019-01-06 125230.7 235.4168 7.5 261.7032 110.3965
## T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 Campaña
## 1 37.69877 2.734000 0.009995935 2.102172 C4
## 2 36.84161 2.580828 0.000000000 1.598925 C4
## 3 38.76250 2.508207 0.315446940 1.854199 C4
## 4 54.19963 2.583115 2.458723773 2.891275 C4
## 5 51.28380 2.082924 2.214033123 2.965717 C4
## 6 52.42363 2.829728 2.652217520 2.966737 C4
## Tonelaje_acumulado_CH1 date_Time Numcampana_med diametro_Mto dias_acum_Ccv
## 1 0.1549793 <NA> <NA> <NA> NA
## 2 0.3019120 <NA> <NA> <NA> NA
## 3 0.4290713 <NA> <NA> <NA> NA
## 4 0.5427367 <NA> <NA> <NA> NA
## 5 0.6578965 <NA> <NA> <NA> NA
## 6 0.7831272 <NA> <NA> <NA> NA
## ton_acum_Ccv pos_med_Ccv remanente_Ccv desgaste_Ccv tasa_desgaste_Ccv_acum
## 1 NA NA <NA> <NA> <NA>
## 2 NA NA <NA> <NA> <NA>
## 3 NA NA <NA> <NA> <NA>
## 4 NA NA <NA> <NA> <NA>
## 5 NA NA <NA> <NA> <NA>
## 6 NA NA <NA> <NA> <NA>
## aplan_Ccv pos_aplan_Ccv
## 1 NA NA
## 2 NA NA
## 3 NA NA
## 4 NA NA
## 5 NA NA
## 6 NA NA
### Estandarizar fechas
datos1$Timestamp <- as.Date(datos1$Timestamp, format = "%Y-%m-%d")
datos1$date_Time <- as.Date(datos1$date_Time, format = "%Y-%m-%d")
### Manipulación de datos y estandarización de los tipos de datos
datos1$Campaña <- gsub("^C", "", datos1$Campaña)
datos1$Numcampana_med <- gsub("^C", "", datos1$Numcampana_med)
### Transformar las columnas a formato numérico reemplazando comas por puntos
datos1$remanente_Ccv <- as.numeric(gsub(",", ".", datos1$remanente_Ccv))
datos1$desgaste_Ccv <- as.numeric(gsub(",", ".", datos1$desgaste_Ccv))
datos1$tasa_desgaste_Ccv_acum <- as.numeric(gsub(",", ".", datos1$tasa_desgaste_Ccv_acum))
### Verificar los cambios
str(datos1[c("remanente_Ccv", "desgaste_Ccv", "tasa_desgaste_Ccv_acum")])
## 'data.frame': 1145 obs. of 3 variables:
## $ remanente_Ccv : num NA NA NA NA NA NA NA NA NA NA ...
## $ desgaste_Ccv : num NA NA NA NA NA NA NA NA NA NA ...
## $ tasa_desgaste_Ccv_acum: num NA NA NA NA NA NA NA NA NA NA ...
Tonelaje_acumulado_CH1: Se reemplazan los valores negativos y 0 por un valor pequeño como 0,01, entendiendo que esto es tonelaje acumulado de avance de campaña y parte en 0 toneladas hacia valores superiores a 15 Millones de toneladas acumuladas en algunos casos.
Presion de hidroset: Se corrigen negativos y los valores 0 se cambian a 0,01 para posibles transformaciones.
T_H: Se corrigen los valores negativos, ya que el tonelaje procesado por dia no puede ser negativo ( 0 se cambia a 0,01).
Corriente: Se corrigen los valores negativos porque el motor consume o no consume corriente esto es igual o mayor a 0 (por futuros analisis el 0 se cambia a 0,01)
“Posion_poste_manto”, “CSS”, “Presion_Hidroset”: valores negativos en instrumentos de medición no pueden ser negativos por la naturaleza de la deteccion y junto con ello los valores igual a 0 tambien se cambian a 0,01 para transformaciones y futuros analisis.
F80_CV_001, F80_CV_701, F80_CV_005 : En el caso de las corres de transporte no tienen valores negativos pero algunos son 0 los cuales se cambian por un valor pequeño que no afecte futuros analisis 0,01.
Las columnas de T_H y Corriente serán actualizadas con los nuevos valores que corregirán los errores captados en las mediciones.
####### la columna añade misma informacion que columna "Numcampana_med"
datos1 <- datos1[, !names(datos1) %in% "Campaña"]
# Definir las columnas a modificar
columnas_a_modificar <- c("Posion_poste_manto", "CSS", "Presion_Hidroset", "F80_CV_001", "F80_CV_701", "F80_CV_005")
# Reemplazar valores <= 0 por 0.01
datos1[columnas_a_modificar] <- lapply(datos1[columnas_a_modificar], function(col) {
col[col <= 0] <- 0.01
return(col)
})
# Reemplazar valores en columnas de interes
indices_T_H <- which(datos1$T_H <= 0)
datos1$T_H[indices_T_H] <- 0.01
indices_Corriente <- which(datos1$Corriente <= 0)
datos1$Corriente[indices_Corriente] <- 0.01
Se observa que el inicio del tonelaje no concordaba con la campaña que estaba en curso, campaña 4_X por lo que se procede a ajustar según el registro del reporte 4_8 con fecha 01/20/2019, junto con ello se observa que la acumulación del tonelaje esta con escala 10^-6.
### Ajuste de escala de Tonelaje_acumulado_CH1 a Toneladas[Ton] x 10^6
datos1$Tonelaje_acumulado_CH1 <- datos1$Tonelaje_acumulado_CH1 * 1000000
datos1$Tonelaje_acumulado_CH1 <- round(datos1$Tonelaje_acumulado_CH1, digits = 0)
# Reemplazar valores iguales o menores a 0 por 0.1 en la columna "Tonelaje_acumulado_CH1"
if ("Tonelaje_acumulado_CH1" %in% names(datos1)) {
datos1$Tonelaje_acumulado_CH1 <- ifelse(datos1$Tonelaje_acumulado_CH1 <= 0, 0.1, datos1$Tonelaje_acumulado_CH1)
}
### Ajustamos tonelajes con la columna T_H
### Asignamos tonelaje obtenido por reporte informacion proporcionada por cliente
datos1$Tonelaje_acumulado_CH1[20:26] <- datos1$ton_acum_Ccv[20:26]
### Recorremos desde fila 19 hasta la 1
if ("Tonelaje_acumulado_CH1" %in% names(datos1) && "T_H" %in% names(datos1)) {
for (i in 19:1) {
datos1$Tonelaje_acumulado_CH1[i] <- datos1$Tonelaje_acumulado_CH1[i + 1] - datos1$T_H[i + 1]
}
} else {
print("Verificar el dataset y el codigo")
}
# Crear una columna para el tonelaje acumulado usando índices
indice_inicio <- 26 # Índice donde empieza la fecha 2019-01-20
indice_final <- 98 # Índice donde termina la fecha 2019-03-21
# Inicializar el tonelaje acumulado en la fecha de inicio
datos1$Tonelaje_acumulado_CH1[indice_inicio] <- datos1$Tonelaje_acumulado_CH1[indice_inicio]
# Calcular el tonelaje acumulado desde el índice de inicio hasta el índice final
for (i in indice_inicio:(indice_final - 1)) {
if (datos1$Timestamp[i] != datos1$Timestamp[i + 1]) {
datos1$Tonelaje_acumulado_CH1[i + 1] <- datos1$Tonelaje_acumulado_CH1[i] + datos1$T_H[i + 1]
} else {
datos1$Tonelaje_acumulado_CH1[i + 1] <- datos1$Tonelaje_acumulado_CH1[i]
}
}
### gráfico para comparar datos de tonelaje acumulado
p1 <- ggplot(data = datos1, aes(x = Timestamp)) +
geom_line(aes(y = Tonelaje_acumulado_CH1, colour = "Tonelaje Acumulado CH1"), size = 1) +
geom_point(aes(y = ton_acum_Ccv, colour = "Tonelaje CCV"), size = 3) +
labs(title = "Comparación de Tonelajes a lo Largo del Tiempo", x = "Fecha", y = "Tonelaje") +
scale_colour_manual(values = c("Tonelaje Acumulado CH1" = "blue", "Tonelaje CCV" = "red")) +
theme_minimal() +
theme(legend.title = element_blank())
## Warning: Using `size` aesthetic for lines was deprecated in ggplot2 3.4.0.
## ℹ Please use `linewidth` instead.
## This warning is displayed once every 8 hours.
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was
## generated.
min_val <- min(c(min(datos1$Tonelaje_acumulado_CH1, na.rm = TRUE), min(datos1$ton_acum_Ccv, na.rm = TRUE)), na.rm = TRUE)
max_val <- max(c(max(datos1$Tonelaje_acumulado_CH1, na.rm = TRUE), max(datos1$ton_acum_Ccv, na.rm = TRUE)), na.rm = TRUE)
p1 <- p1 + scale_y_continuous(limits = c(min_val, max_val * 1.1))
print(p1)
## Warning: Removed 921 rows containing missing values or values outside the scale range
## (`geom_point()`).
estadisticas1 <- datos1 %>%
summarise(
Media_Tonelaje = mean(Tonelaje_acumulado_CH1, na.rm = TRUE),
Mediana_Tonelaje = median(Tonelaje_acumulado_CH1, na.rm = TRUE),
Moda_Tonelaje = as.numeric(names(which.max(table(Tonelaje_acumulado_CH1)))),
Max_Tonelaje = max(Tonelaje_acumulado_CH1, na.rm = TRUE),
Min_Tonelaje = min(Tonelaje_acumulado_CH1, na.rm = TRUE),
Media_CCV = mean(ton_acum_Ccv, na.rm = TRUE),
Mediana_CCV = median(ton_acum_Ccv, na.rm = TRUE),
Moda_CCV = as.numeric(names(which.max(table(ton_acum_Ccv)))),
Max_CCV = max(ton_acum_Ccv, na.rm = TRUE),
Min_CCV = min(ton_acum_Ccv, na.rm = TRUE)
)
kable(estadisticas1, format = "markdown", caption = "Estadísticas Descriptivas de Tonelaje")
| Media_Tonelaje | Mediana_Tonelaje | Moda_Tonelaje | Max_Tonelaje | Min_Tonelaje | Media_CCV | Mediana_CCV | Moda_CCV | Max_CCV | Min_CCV |
|---|---|---|---|---|---|---|---|---|---|
| 10845232 | 11227616 | 0.1 | 23172738 | 0.1 | 10307880 | 10623255 | 0 | 22677104 | 0 |
correlacion <- cor(datos1$Tonelaje_acumulado_CH1, datos1$ton_acum_Ccv, use = "complete.obs")
print(paste("La correlación entre Tonelaje_acumulado_CH1 y ton_acum_Ccv es:", correlacion))
## [1] "La correlación entre Tonelaje_acumulado_CH1 y ton_acum_Ccv es: 0.999128414782128"
# Crear una nueva variable para distinguir los puntos en el gráfico
datos1$color_var <- ifelse(datos1$Tonelaje_acumulado_CH1 > datos1$ton_acum_Ccv, "Tonelaje CH1 mayor", "Tonelaje CCV mayor")
# Filtrar filas donde color_var no es NA
datos1_filtrado <- subset(datos1, !is.na(color_var))
ggplot(datos1_filtrado, aes(x = Tonelaje_acumulado_CH1, y = ton_acum_Ccv, color = color_var)) +
geom_point(alpha = 0.6) + # Ajustar la transparencia de los puntos
labs(title = "Gráfico de Dispersión entre Tonelaje Acumulado CH1 y Tonelaje CCV",
x = "Tonelaje Acumulado CH1",
y = "Tonelaje CCV") +
scale_color_manual(values = c("Tonelaje CH1 mayor" = "aquamarine", "Tonelaje CCV mayor" = "lightcoral")) +
theme_minimal()
# Elimina color_var
datos1$color_var <- NULL
# Selecciona las columnas que deseas incluir en la matriz de correlación
# Aquí incluyo Tonelaje_acumulado_CH1 y ton_acum_Ccv como ejemplo; añade más columnas si es necesario
correlacion9 <- cor(datos1[, c("Tonelaje_acumulado_CH1", "ton_acum_Ccv")], use = "complete.obs")
# Verifica la matriz de correlación
print(correlacion9)
## Tonelaje_acumulado_CH1 ton_acum_Ccv
## Tonelaje_acumulado_CH1 1.0000000 0.9991284
## ton_acum_Ccv 0.9991284 1.0000000
Al observar la comparación de Tonelaje_acumulado_CH1 y ton_acum_Ccv en el gráfico de serie de tiempo, comparar valores estadisticos y la alta correlación (0.9) se determina eliminar la columna ton_acum_Ccv, entendiendo que la naturaleza de los datos sensorizados permite tener una mayor presición de los datos y con el objetivo de ahorrar costos y de tener una columna que podría inducir a pequeños errores ya que en la mayoria de los puntos existe un delta al comparar las columnas de interes, esto ocurre principalmente porque los tonelajes que se indican en los reportes son los proporcionados por cliente el cual podría tener diferencia con lo real (Datos sensorizados).
Datos Operacionales: Estos datos se registran
diariamente y proporcionan una visión continua del funcionamiento del
chancador.
Datos de Reportes: Estos datos se capturan únicamente
en fechas específicas, generalmente cuando el chancador está detenido
por mantenimiento. Por lo tanto, no tienen la misma continuidad que los
datos operacionales. Se procede a dividir los conjuntos de datos.
Dado que los datos de mantenimiento no están presentes todos los días, se decide dividir el conjunto de datos en dos: Este enfoque facilitará el análisis específico en cada tipo de datos y evitará confusiones en la imputación de datos faltantes, que en el caso de los datos de reportes no son realmente “faltantes” sino “no aplicables” para días sin mantenimiento.
### Se elimina ton_acum_ccv
datos1 <- subset(datos1, select = -ton_acum_Ccv)
# Datos de reportes (excluyendo columnas de datos operacionales)
datosRepo <- datos1[ , !(names(datos1) %in% c("Timestamp", "T_H", "Posion_poste_manto", "CSS", "Presion_Hidroset", "Corriente", "T_aceite_motor", "F80_CV_001", "F80_CV_701", "F80_CV_005", "Campaña", "Tonelaje_acumulado_CH1"))]
# Datos operacionales (incluyendo solo columnas operacionales)
datosOpe <- datos1[ , c("Timestamp", "T_H", "Posion_poste_manto", "CSS", "Presion_Hidroset", "Corriente", "T_aceite_motor", "F80_CV_001", "F80_CV_701", "F80_CV_005", "Tonelaje_acumulado_CH1")]
###Información
summary(datosOpe)
## Timestamp T_H Posion_poste_manto CSS
## Min. :2019-01-01 Min. : 0.01 Min. : 0.01 Min. : 0.010
## 1st Qu.:2019-08-27 1st Qu.: 58252.83 1st Qu.:105.30 1st Qu.: 7.500
## Median :2020-05-03 Median : 85779.60 Median :152.85 Median : 7.500
## Mean :2020-04-24 Mean : 75268.35 Mean :155.73 Mean : 7.643
## 3rd Qu.:2020-12-15 3rd Qu.: 99023.01 3rd Qu.:212.91 3rd Qu.: 7.729
## Max. :2021-08-10 Max. :169829.45 Max. :263.73 Max. :69.375
##
## Presion_Hidroset Corriente T_aceite_motor F80_CV_001
## Min. : 0.01 Min. : 0.01 Min. : 4.331 Min. :0.001133
## 1st Qu.:227.93 1st Qu.: 77.99 1st Qu.:47.916 1st Qu.:1.898283
## Median :249.08 Median : 99.86 Median :52.740 Median :2.779308
## Mean :238.19 Mean : 88.20 Mean :48.568 Mean :2.506096
## 3rd Qu.:270.36 3rd Qu.:110.25 3rd Qu.:55.047 3rd Qu.:3.313581
## Max. :328.38 Max. :141.98 Max. :58.916 Max. :5.659055
## NA's :1
## F80_CV_701 F80_CV_005 Tonelaje_acumulado_CH1
## Min. :0.001308 Min. :0.010 Min. : 0
## 1st Qu.:1.842345 1st Qu.:2.262 1st Qu.: 5056156
## Median :2.602227 Median :2.918 Median :11227616
## Mean :2.385522 Mean :2.775 Mean :10845232
## 3rd Qu.:3.154039 3rd Qu.:3.455 3rd Qu.:16589227
## Max. :5.080948 Max. :5.998 Max. :23172738
##
head(datosOpe)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset Corriente
## 1 2019-01-01 154979.3 202.2061 7.5 258.3561 102.9613
## 2 2019-01-02 146932.7 204.2289 7.5 251.4526 102.3415
## 3 2019-01-03 127159.3 193.3790 7.5 244.4037 101.0957
## 4 2019-01-04 113665.4 217.8944 7.5 248.4902 103.3990
## 5 2019-01-05 115159.8 227.8408 7.5 241.4058 101.6643
## 6 2019-01-06 125230.7 235.4168 7.5 261.7032 110.3965
## T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 Tonelaje_acumulado_CH1
## 1 37.69877 2.734000 0.009995935 2.102172 11437411
## 2 36.84161 2.580828 0.010000000 1.598925 11584344
## 3 38.76250 2.508207 0.315446940 1.854199 11711503
## 4 54.19963 2.583115 2.458723773 2.891275 11825169
## 5 51.28380 2.082924 2.214033123 2.965717 11940328
## 6 52.42363 2.829728 2.652217520 2.966737 12065559
str(datosOpe)
## 'data.frame': 1145 obs. of 11 variables:
## $ Timestamp : Date, format: "2019-01-01" "2019-01-02" ...
## $ T_H : num 154979 146933 127159 113665 115160 ...
## $ Posion_poste_manto : num 202 204 193 218 228 ...
## $ CSS : num 7.5 7.5 7.5 7.5 7.5 ...
## $ Presion_Hidroset : num 258 251 244 248 241 ...
## $ Corriente : num 103 102 101 103 102 ...
## $ T_aceite_motor : num 37.7 36.8 38.8 54.2 51.3 ...
## $ F80_CV_001 : num 2.73 2.58 2.51 2.58 2.08 ...
## $ F80_CV_701 : num 0.01 0.01 0.315 2.459 2.214 ...
## $ F80_CV_005 : num 2.1 1.6 1.85 2.89 2.97 ...
## $ Tonelaje_acumulado_CH1: num 11437411 11584344 11711503 11825169 11940328 ...
summary(datosRepo)
## date_Time Numcampana_med diametro_Mto dias_acum_Ccv
## Min. :2019-01-20 Length:1145 Length:1145 Min. : 0.0
## 1st Qu.:2019-09-02 Class :character Class :character 1st Qu.: 51.5
## Median :2020-06-08 Mode :character Mode :character Median :125.5
## Mean :2020-05-11 Mean :120.4
## 3rd Qu.:2020-12-19 3rd Qu.:180.8
## Max. :2021-08-10 Max. :280.0
## NA's :921 NA's :921
## pos_med_Ccv remanente_Ccv desgaste_Ccv tasa_desgaste_Ccv_acum
## Min. :1 Min. : 43.40 Min. : 0.00 Min. :0.000
## 1st Qu.:2 1st Qu.: 76.55 1st Qu.: 15.15 1st Qu.:3.100
## Median :4 Median :105.10 Median : 39.00 Median :3.800
## Mean :4 Mean :105.56 Mean : 39.99 Mean :3.369
## 3rd Qu.:6 3rd Qu.:128.90 3rd Qu.: 63.50 3rd Qu.:4.300
## Max. :7 Max. :188.90 Max. :101.50 Max. :5.200
## NA's :921 NA's :921 NA's :921 NA's :921
## aplan_Ccv pos_aplan_Ccv
## Min. :-15.300 Min. : 0
## 1st Qu.: -7.300 1st Qu.:100
## Median : -4.750 Median :300
## Mean : -3.794 Mean :300
## 3rd Qu.: 0.100 3rd Qu.:500
## Max. : 6.000 Max. :600
## NA's :921 NA's :921
head(datosRepo)
## date_Time Numcampana_med diametro_Mto dias_acum_Ccv pos_med_Ccv remanente_Ccv
## 1 <NA> <NA> <NA> NA NA NA
## 2 <NA> <NA> <NA> NA NA NA
## 3 <NA> <NA> <NA> NA NA NA
## 4 <NA> <NA> <NA> NA NA NA
## 5 <NA> <NA> <NA> NA NA NA
## 6 <NA> <NA> <NA> NA NA NA
## desgaste_Ccv tasa_desgaste_Ccv_acum aplan_Ccv pos_aplan_Ccv
## 1 NA NA NA NA
## 2 NA NA NA NA
## 3 NA NA NA NA
## 4 NA NA NA NA
## 5 NA NA NA NA
## 6 NA NA NA NA
str(datosRepo)
## 'data.frame': 1145 obs. of 10 variables:
## $ date_Time : Date, format: NA NA ...
## $ Numcampana_med : chr NA NA NA NA ...
## $ diametro_Mto : chr NA NA NA NA ...
## $ dias_acum_Ccv : num NA NA NA NA NA NA NA NA NA NA ...
## $ pos_med_Ccv : num NA NA NA NA NA NA NA NA NA NA ...
## $ remanente_Ccv : num NA NA NA NA NA NA NA NA NA NA ...
## $ desgaste_Ccv : num NA NA NA NA NA NA NA NA NA NA ...
## $ tasa_desgaste_Ccv_acum: num NA NA NA NA NA NA NA NA NA NA ...
## $ aplan_Ccv : num NA NA NA NA NA NA NA NA NA NA ...
## $ pos_aplan_Ccv : num NA NA NA NA NA NA NA NA NA NA ...
#### Eliminar las fechas repetidas en dataset datosOpe (manipulación al unir bases de datos)
datosOpe <- datosOpe %>%
distinct(Timestamp, .keep_all = TRUE)
# Filtrar las filas que contienen datos completos en datosRepo
datosRepo <- datosRepo[complete.cases(datosRepo), ]
row.names(datosRepo) <- NULL
head(datosRepo)
## date_Time Numcampana_med diametro_Mto dias_acum_Ccv pos_med_Ccv
## 1 2019-01-20 4_8 115 149 7
## 2 2019-01-20 4_8 115 149 6
## 3 2019-01-20 4_8 115 149 5
## 4 2019-01-20 4_8 115 149 4
## 5 2019-01-20 4_8 115 149 3
## 6 2019-01-20 4_8 115 149 2
## remanente_Ccv desgaste_Ccv tasa_desgaste_Ccv_acum aplan_Ccv pos_aplan_Ccv
## 1 71.1 43.8 3.4 3.7 600
## 2 88.4 47.2 3.7 0.7 500
## 3 80.3 56.4 4.4 -2.6 400
## 4 78.3 62.7 4.9 -4.9 300
## 5 80.0 66.3 5.2 -7.2 200
## 6 92.0 63.5 4.9 -8.1 100
# Contar valores faltantes en datos de operacion
countNa1 <- sapply(datosOpe, function(x) sum(is.na(x)))
kable(data.frame(Valores_Faltantes = countNa1), format = "markdown")
| Valores_Faltantes | |
|---|---|
| Timestamp | 0 |
| T_H | 0 |
| Posion_poste_manto | 0 |
| CSS | 0 |
| Presion_Hidroset | 0 |
| Corriente | 1 |
| T_aceite_motor | 0 |
| F80_CV_001 | 0 |
| F80_CV_701 | 0 |
| F80_CV_005 | 0 |
| Tonelaje_acumulado_CH1 | 0 |
# Contar valores faltantes en datos de reportes
countNa2 <- sapply(datosRepo, function(x) sum(is.na(x)))
kable(data.frame(Valores_Faltantes = countNa2), format = "markdown")
| Valores_Faltantes | |
|---|---|
| date_Time | 0 |
| Numcampana_med | 0 |
| diametro_Mto | 0 |
| dias_acum_Ccv | 0 |
| pos_med_Ccv | 0 |
| remanente_Ccv | 0 |
| desgaste_Ccv | 0 |
| tasa_desgaste_Ccv_acum | 0 |
| aplan_Ccv | 0 |
| pos_aplan_Ccv | 0 |
### Obtener indice de la fila, faltante de datosOpe
filasNa <- function(df) {
indiceNa <- which(is.na(df), arr.ind = TRUE)
df_na <- data.frame(Fila = indiceNa[, 1], Columna = colnames(df)[indiceNa[, 2]])
return(df_na)
}
RepoNa2 <- filasNa(datosOpe)
kable(RepoNa2, format = "markdown")
| Fila | Columna | |
|---|---|---|
| row | 884 | Corriente |
### Extraer la fila 884
fila_884 <- datosOpe[884, , drop = FALSE]
kable(fila_884, caption = "fila 884")
| Timestamp | T_H | Posion_poste_manto | CSS | Presion_Hidroset | Corriente | T_aceite_motor | F80_CV_001 | F80_CV_701 | F80_CV_005 | Tonelaje_acumulado_CH1 | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 884 | 2021-06-02 | 0.01 | 101.2008 | 7.5 | 190.3523 | NA | 55.87287 | 0.01 | 2.764365 | 3.561787 | 803349 |
Se observa que el dato faltante de corriente es un día en que el tonelaje procesado es -198,5993 [A] lo cual indica que ese día no se captaron datos de que el chancador estaba operativo. Esta condición puede ser debido a una falta de los datos sensorizados por problemas puntuales o tal vez una parada no programada durante el día. Para el tratamiento de este valor faltante se utilizará como referencia la columna de tonelaje procesado diario (T_H)
Para el valor faltante en corriente se realiza una imputación basada en modelos
# Verificar y visualizar valores faltantes
sum(is.na(datosOpe$Corriente))
## [1] 1
# Visualizar la relación entre T_H y Corriente
ggplot(datosOpe, aes(x = T_H, y = Corriente)) +
geom_point() +
# geom_smooth(method = "lm", se = TRUE, color = "blue") +
labs(title = "Scatterplot de T_H vs Corriente",
x = "Tonelaje procesado diario (T_H)", y = "Corriente")
## Warning: Removed 1 row containing missing values or values outside the scale range
## (`geom_point()`).
# Preparar datos para el modelo de regresión, excluyendo NA
datosOpe_clean <- na.omit(datosOpe)
# Crear un modelo de regresión lineal
model <- lm(Corriente ~ T_H, data = datosOpe_clean)
# Identificar puntos con residuos estandarizados mayores a 2
model_diag <- augment(model)
high_resid_points <- model_diag %>%
filter(abs(.std.resid) > 2)
# Visualizar gráficamente los puntos con residuos altos
ggplot(datosOpe_clean, aes(x = T_H, y = Corriente)) +
geom_point() +
geom_smooth(method = "lm", col = "blue", fill = "blue", level = 0.95) +
geom_text(data = high_resid_points, aes(label = sprintf("%.2f", Corriente)),
nudge_y = 1, check_overlap = TRUE, color = "red") +
labs(title = "Relación entre T_H y Corriente con Residuos Altos",
x = "T_H", y = "Corriente")
## `geom_smooth()` using formula = 'y ~ x'
# Predecir y añadir el valor faltante
missing_index <- which(is.na(datosOpe$Corriente))
predicted_corriente <- predict(model, newdata = datosOpe[missing_index, , drop = FALSE])
datosOpe$Corriente[missing_index] <- predicted_corriente
conteo_valores_no_positivos <- datosOpe %>%
select(-Timestamp) %>%
summarise(across(everything(), ~ sum(. <= 0)))
# Convertir los resultados a una tabla formateada usando kable
tabla_valores_no_positivos <- conteo_valores_no_positivos %>%
pivot_longer(cols = everything(), names_to = "Columna", values_to = "Valores_no_positivos") %>%
kable(format = "markdown", col.names = c("Columna", "Valores No Positivos"))
# Mostrar la tabla
print(tabla_valores_no_positivos)
##
##
## |Columna | Valores No Positivos|
## |:----------------------|--------------------:|
## |T_H | 0|
## |Posion_poste_manto | 0|
## |CSS | 0|
## |Presion_Hidroset | 0|
## |Corriente | 0|
## |T_aceite_motor | 0|
## |F80_CV_001 | 0|
## |F80_CV_701 | 0|
## |F80_CV_005 | 0|
## |Tonelaje_acumulado_CH1 | 0|
datosOpe <- datosOpe %>% rename(ton_acum = Tonelaje_acumulado_CH1)
datosOpe1 <- datosOpe
datosRepo1 <- datosRepo
datosOpe1 <- datosOpe1[, c("Timestamp", "T_H", "Posion_poste_manto", "CSS",
"Presion_Hidroset", "Corriente", "T_aceite_motor",
"F80_CV_001", "F80_CV_701", "F80_CV_005", "ton_acum")]
# Seleccionar columnas numéricas
numeric_cols <- sapply(datosOpe1, is.numeric)
numeric_data <- datosOpe1[, numeric_cols]
# Visualización de outliers utilizando boxplots
# Crear una lista vacía para almacenar gráficos
plots <- list()
# Generar un boxplot para cada columna numérica
for (col in colnames(numeric_data)) {
p <- ggplot(datosOpe1, aes(y = !!sym(col))) +
geom_boxplot() +
ggtitle(paste("Boxplot of", col))
plots[[col]] <- p
}
# Mostrar los primeros 6 gráficos en una grilla de 2x3
do.call("grid.arrange", c(plots[1:6], ncol = 2, nrow = 3))
# Mostrar los siguientes gráficos en otra grilla de 2x3
if (length(plots) > 6) {
do.call("grid.arrange", c(plots[7:12], ncol = 2, nrow = 3))
}
# Función para detectar outliers usando IQR
detect_outliers <- function(df, col) {
Q1 <- quantile(df[[col]], 0.25, na.rm = TRUE)
Q3 <- quantile(df[[col]], 0.75, na.rm = TRUE)
IQR <- Q3 - Q1
outliers <- which(df[[col]] < (Q1 - 1.5 * IQR) | df[[col]] > (Q3 + 1.5 * IQR))
return(outliers)
}
# Detectar outliers para cada columna numérica y asignar a variables específicas
T_H_outliers <- datosOpe1[detect_outliers(numeric_data, "T_H"), ]
Posion_poste_manto_outliers <- datosOpe1[detect_outliers(numeric_data, "Posion_poste_manto"), ]
CSS_outliers <- datosOpe1[detect_outliers(numeric_data, "CSS"), ]
Presion_Hidroset_outliers <- datosOpe1[detect_outliers(numeric_data, "Presion_Hidroset"), ]
Corriente_outliers <- datosOpe1[detect_outliers(numeric_data, "Corriente"), ]
T_aceite_motor_outliers <- datosOpe1[detect_outliers(numeric_data, "T_aceite_motor"), ]
F80_CV_001_outliers <- datosOpe1[detect_outliers(numeric_data, "F80_CV_001"), ]
F80_CV_701_outliers <- datosOpe1[detect_outliers(numeric_data, "F80_CV_701"), ]
F80_CV_005_outliers <- datosOpe1[detect_outliers(numeric_data, "F80_CV_005"), ]
ton_acum_outliers <- datosOpe1[detect_outliers(numeric_data, "ton_acum"), ]
# Crear un dataframe con el conteo de outliers
outlier_counts <- data.frame(
Variable = c("T_H", "Posion_poste_manto", "CSS", "Presion_Hidroset",
"Corriente", "T_aceite_motor", "F80_CV_001",
"F80_CV_701", "F80_CV_005", "ton_acum"),
Outliers = c(nrow(T_H_outliers), nrow(Posion_poste_manto_outliers),
nrow(CSS_outliers), nrow(Presion_Hidroset_outliers), nrow(Corriente_outliers),
nrow(T_aceite_motor_outliers), nrow(F80_CV_001_outliers),
nrow(F80_CV_701_outliers), nrow(F80_CV_005_outliers),
nrow(ton_acum_outliers))
)
# Mostrar la tabla de conteo de outliers con kable
kable(outlier_counts, caption = "Conteo de Outliers por Variable")
| Variable | Outliers |
|---|---|
| T_H | 85 |
| Posion_poste_manto | 0 |
| CSS | 19 |
| Presion_Hidroset | 44 |
| Corriente | 102 |
| T_aceite_motor | 111 |
| F80_CV_001 | 66 |
| F80_CV_701 | 65 |
| F80_CV_005 | 51 |
| ton_acum | 0 |
head(datosOpe1, 30)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset Corriente
## 1 2019-01-01 154979.34 202.20612 7.500000 258.356143 102.96132
## 2 2019-01-02 146932.65 204.22895 7.500000 251.452566 102.34145
## 3 2019-01-03 127159.29 193.37900 7.500000 244.403748 101.09568
## 4 2019-01-04 113665.43 217.89443 7.500000 248.490238 103.39896
## 5 2019-01-05 115159.79 227.84076 7.500000 241.405772 101.66433
## 6 2019-01-06 125230.67 235.41683 7.500000 261.703197 110.39655
## 7 2019-01-07 115360.01 232.59723 7.541667 284.844956 123.82054
## 8 2019-01-08 94863.25 228.05275 8.000000 267.629224 112.82917
## 9 2019-01-09 87410.02 231.96556 8.000000 265.963168 116.76076
## 10 2019-01-10 95211.22 244.15895 8.000000 276.401559 125.12916
## 11 2019-01-11 90695.03 243.95059 8.000000 262.114219 110.19245
## 12 2019-01-12 99373.24 244.06586 8.000000 277.813777 124.49621
## 13 2019-01-13 103179.08 245.58860 8.000000 246.049255 102.54717
## 14 2019-01-14 46200.84 247.94470 8.000000 219.257179 78.02351
## 15 2019-01-15 0.01 225.93630 7.958333 135.151396 0.01000
## 16 2019-01-16 0.01 202.84693 7.541667 8.556933 0.01000
## 17 2019-01-17 0.01 202.03400 8.000000 8.611637 0.01000
## 18 2019-01-18 0.01 201.25147 7.958333 8.742568 0.01000
## 19 2019-01-19 0.01 53.47405 7.500000 157.181469 0.01000
## 20 2019-01-20 36531.10 42.34228 7.500000 229.999434 46.61913
## 21 2019-01-21 92685.52 53.29738 7.500000 282.245209 105.18909
## 22 2019-01-22 87745.61 53.46137 7.500000 273.604721 110.12873
## 23 2019-01-23 56735.67 53.52586 7.500000 253.193405 95.34219
## 24 2019-01-24 83225.47 59.32192 7.500000 256.770877 100.89891
## 25 2019-01-25 88530.11 64.27847 7.500000 259.253080 101.38239
## 26 2019-01-26 80042.94 64.25159 7.500000 264.936851 93.39901
## 27 2019-01-27 102120.51 63.01171 7.500000 262.761858 105.45590
## 28 2019-01-28 102124.02 64.02501 7.500000 261.996204 106.33217
## 29 2019-01-29 97324.06 65.24362 7.500000 281.117611 109.79523
## 30 2019-01-30 96546.87 72.13519 7.500000 275.134096 105.26992
## T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 ton_acum
## 1 37.69877 2.734000 0.009995935 2.1021723 11437411
## 2 36.84161 2.580828 0.010000000 1.5989251 11584344
## 3 38.76250 2.508207 0.315446940 1.8541986 11711503
## 4 54.19963 2.583115 2.458723773 2.8912753 11825169
## 5 51.28380 2.082924 2.214033123 2.9657168 11940328
## 6 52.42363 2.829728 2.652217520 2.9667371 12065559
## 7 51.24578 3.847832 3.032121758 3.9391700 12180919
## 8 44.96757 3.327394 1.301360822 2.7061587 12275782
## 9 46.03698 3.421496 1.401532981 3.1142770 12363192
## 10 47.23661 3.813759 1.417639070 3.2235497 12458404
## 11 50.38093 2.885906 2.956043905 3.5066171 12549099
## 12 51.17236 3.763040 2.996162328 3.7315247 12648472
## 13 50.06650 2.595243 2.576847181 3.1771571 12751651
## 14 52.27089 1.206757 3.228821277 2.9989041 12797852
## 15 40.65341 0.010000 0.088466645 0.2174042 12797852
## 16 24.99891 0.010000 0.010000000 0.0100000 12797852
## 17 22.53622 0.010000 0.010000000 0.0100000 12797852
## 18 18.33353 0.010000 0.010000000 0.0100000 12797852
## 19 19.00045 0.010000 0.010000000 0.0100000 12797852
## 20 46.89799 5.659055 1.705099911 2.2054545 12834383
## 21 54.42911 3.338046 3.759763271 3.0435907 12927069
## 22 52.07658 2.960537 2.743710732 2.6977181 13014814
## 23 48.53875 1.942956 1.873627737 2.6085556 13071550
## 24 48.89649 2.345388 2.613457567 2.4108110 13154775
## 25 49.38316 2.570862 3.062152814 2.7436641 13243305
## 26 51.85536 2.572981 3.608127311 2.6547584 13323348
## 27 48.81553 2.601890 2.892583454 2.9360754 13425469
## 28 48.31176 2.755972 2.841873056 3.1307126 13527593
## 29 45.49663 3.458478 1.741398185 2.7730836 13624917
## 30 50.57423 3.129912 2.438373041 3.4628318 13721464
head(T_H_outliers, 30)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset
## 1 2019-01-01 154979.34 202.20612 7.500000 258.356143
## 2 2019-01-02 146932.65 204.22895 7.500000 251.452566
## 15 2019-01-15 0.01 225.93630 7.958333 135.151396
## 16 2019-01-16 0.01 202.84693 7.541667 8.556933
## 17 2019-01-17 0.01 202.03400 8.000000 8.611637
## 18 2019-01-18 0.01 201.25147 7.958333 8.742568
## 19 2019-01-19 0.01 53.47405 7.500000 157.181469
## 20 2019-01-20 36531.10 42.34228 7.500000 229.999434
## 43 2019-02-12 19958.07 112.94737 69.375000 222.760463
## 44 2019-02-13 19947.83 114.82612 7.500000 218.956877
## 77 2019-03-18 21195.62 198.49698 7.500000 219.192808
## 78 2019-03-19 0.01 22.05181 7.500000 34.429735
## 79 2019-03-20 0.01 0.01000 7.500000 0.010000
## 80 2019-03-21 0.01 0.01000 7.500000 0.010000
## 81 2019-03-22 0.01 18.59956 7.500000 31.421540
## 127 2019-05-07 27973.27 200.55711 7.500000 216.577735
## 149 2019-05-29 0.01 55.61887 7.500000 88.634211
## 150 2019-05-30 22511.35 53.99816 7.500000 209.187638
## 175 2019-06-24 19139.13 70.17309 7.500000 207.580697
## 176 2019-06-25 0.01 58.48542 7.500000 119.184292
## 177 2019-06-26 0.01 47.42849 7.500000 7.871639
## 178 2019-06-27 0.01 46.38774 7.500000 117.261253
## 179 2019-06-28 0.01 43.95892 7.500000 175.376969
## 180 2019-06-29 0.01 43.20602 7.500000 177.899915
## 226 2019-08-14 26770.06 207.58160 7.500000 219.750329
## 253 2019-09-10 30311.12 245.51034 7.750000 217.883641
## 254 2019-09-11 0.01 21.31455 7.750000 64.219840
## 255 2019-09-12 0.01 20.44627 7.729167 40.623695
## 281 2019-10-08 31308.55 125.87614 7.500000 228.451514
## 282 2019-10-09 20348.15 126.18405 7.500000 211.075542
## Corriente T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 ton_acum
## 1 102.9613236 37.69877 2.733999724 0.009995935 2.1021723 11437411.3
## 2 102.3414507 36.84161 2.580828490 0.010000000 1.5989251 11584344.0
## 15 0.0100000 40.65341 0.010000000 0.088466645 0.2174042 12797851.9
## 16 0.0100000 24.99891 0.010000000 0.010000000 0.0100000 12797851.9
## 17 0.0100000 22.53622 0.010000000 0.010000000 0.0100000 12797851.9
## 18 0.0100000 18.33353 0.010000000 0.010000000 0.0100000 12797851.9
## 19 0.0100000 19.00045 0.010000000 0.010000000 0.0100000 12797851.9
## 20 46.6191311 46.89799 5.659054995 1.705099911 2.2054545 12834383.0
## 43 51.7660758 37.33947 1.080412596 0.817207394 0.8285527 14978014.0
## 44 61.3042102 44.57258 0.735764346 1.904086001 1.5382229 14997961.8
## 77 53.7778369 52.24535 0.844764545 3.006839651 2.0315111 17967475.8
## 78 0.0619501 42.49164 0.012151670 1.547525644 1.1348397 17967475.8
## 79 0.0100000 38.93675 0.010000000 1.728804133 1.2295940 17967475.8
## 80 0.0100000 53.99001 0.010000000 4.304394859 4.5101251 17967475.8
## 81 0.9404026 51.22423 0.010000000 3.565637715 2.8647831 0.1
## 127 51.8322726 52.34146 1.219738799 2.212686751 1.9522210 4026882.0
## 149 0.0100000 35.24328 0.011239038 1.570499680 1.1073880 5768427.0
## 150 33.7159783 46.10239 0.682916357 2.983121260 2.1345206 5790938.0
## 175 43.5276566 35.97505 0.650703190 0.586644136 2.8296758 8020195.0
## 176 0.0100000 30.35013 0.010000000 0.010000000 2.1450566 8019928.0
## 177 0.0100000 20.65058 0.010000000 0.010000000 0.1059562 8019683.0
## 178 0.0100000 10.69983 0.010000000 0.010000000 0.0100000 8019451.0
## 179 0.0100000 18.61315 0.010000000 0.010000000 0.0100000 8019190.0
## 180 0.0100000 22.92254 0.010000000 0.010000000 0.0100000 8018844.0
## 226 42.3421670 42.14589 0.976805818 1.096268172 1.1339708 12267329.0
## 253 50.7894257 39.31355 0.991761042 1.186440272 1.3296140 14680440.0
## 254 0.0100000 20.06790 0.010000000 0.010000000 0.0100000 14680160.0
## 255 0.5496111 52.08510 0.002948072 2.494095810 2.1395124 14679944.0
## 281 49.6209374 38.83861 1.106780946 1.535402492 1.3535701 17039069.0
## 282 49.1702896 38.51890 0.513771874 0.827849625 0.9192092 17059417.0
head(Posion_poste_manto_outliers, 30)
## [1] Timestamp T_H Posion_poste_manto CSS
## [5] Presion_Hidroset Corriente T_aceite_motor F80_CV_001
## [9] F80_CV_701 F80_CV_005 ton_acum
## <0 rows> (o 0- extensión row.names)
head(CSS_outliers, 30)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset
## 42 2019-02-11 84297.28 118.931419 13.1250000 269.615781
## 43 2019-02-12 19958.07 112.947370 69.3750000 222.760463
## 70 2019-03-11 79722.74 192.960725 13.1250000 247.437310
## 71 2019-03-12 98938.74 195.357770 69.3750000 255.008110
## 92 2019-04-02 97316.76 106.991066 13.1250000 249.685680
## 93 2019-04-03 63530.25 107.972879 69.3750000 216.295841
## 97 2019-04-07 86893.70 128.604477 10.3125000 242.015924
## 336 2019-12-02 29696.15 206.978751 6.8750000 220.744177
## 337 2019-12-03 0.01 216.156604 0.0100000 68.632945
## 338 2019-12-04 0.01 201.157492 0.0100000 7.432575
## 339 2019-12-05 0.01 6.402414 0.6666667 16.533794
## 619 2020-09-10 46597.52 204.900767 7.1041667 224.176504
## 620 2020-09-11 0.01 110.407927 0.6458333 79.527271
## 621 2020-09-12 0.01 24.409211 7.1041667 8.493155
## 622 2020-09-13 0.01 24.083224 0.6250000 8.233641
## 870 2021-05-19 0.01 0.010000 0.0100000 0.010000
## 871 2021-05-20 0.01 0.010000 0.0100000 0.010000
## 872 2021-05-21 0.01 0.010000 0.0100000 0.010000
## 873 2021-05-22 0.01 45.620566 0.3333333 65.958992
## Corriente T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 ton_acum
## 42 93.35662 43.682505 2.485510186 1.8100015 2.3286020 14958055.9
## 43 51.76608 37.339466 1.080412596 0.8172074 0.8285527 14978014.0
## 70 95.54054 49.204710 1.922645767 2.3007710 2.5750057 17341369.9
## 71 89.06010 41.780800 2.300869038 0.6124986 2.1671045 17440308.7
## 92 104.36324 49.960675 3.147430157 2.0954293 3.1071506 1079672.0
## 93 64.33940 48.589427 1.418567290 2.2283462 2.2736836 1143202.0
## 97 100.26219 49.988680 2.741547033 2.1523903 2.5574918 1484209.0
## 336 44.71675 55.880085 0.803888843 3.0916569 3.0095417 21215300.0
## 337 0.01000 56.511999 0.010000000 3.6901907 2.7138747 21215100.0
## 338 0.01000 57.343332 0.010000000 3.5076268 2.1995731 21214882.0
## 339 0.01000 55.680846 0.010000000 3.5469368 3.2539460 0.1
## 619 55.13612 44.960570 1.179482848 1.1370821 1.7160282 23172738.0
## 620 0.01000 11.669826 0.056326417 0.0100000 0.0100000 23172675.0
## 621 0.01000 8.432544 0.010000000 0.0100000 0.0100000 23172477.0
## 622 0.01000 29.552613 0.010000000 1.2205839 0.9432721 23172256.0
## 870 0.01000 56.034353 0.003108983 2.2862278 2.8211925 21967260.0
## 871 0.01000 55.057588 0.010000000 2.7128299 3.3762369 21966980.0
## 872 0.01000 35.045632 0.010000000 1.0460979 2.6827996 21966714.0
## 873 0.01000 4.428674 0.010000000 0.0100000 0.5514292 0.1
head(Presion_Hidroset_outliers, 30)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset
## 15 2019-01-15 0.01 225.936304 7.9583333 135.151396
## 16 2019-01-16 0.01 202.846935 7.5416667 8.556933
## 17 2019-01-17 0.01 202.034004 8.0000000 8.611637
## 18 2019-01-18 0.01 201.251470 7.9583333 8.742568
## 19 2019-01-19 0.01 53.474046 7.5000000 157.181469
## 78 2019-03-19 0.01 22.051813 7.5000000 34.429735
## 79 2019-03-20 0.01 0.010000 7.5000000 0.010000
## 80 2019-03-21 0.01 0.010000 7.5000000 0.010000
## 81 2019-03-22 0.01 18.599555 7.5000000 31.421540
## 149 2019-05-29 0.01 55.618873 7.5000000 88.634211
## 176 2019-06-25 0.01 58.485421 7.5000000 119.184292
## 177 2019-06-26 0.01 47.428494 7.5000000 7.871639
## 178 2019-06-27 0.01 46.387738 7.5000000 117.261253
## 179 2019-06-28 0.01 43.958922 7.5000000 175.376969
## 180 2019-06-29 0.01 43.206022 7.5000000 177.899915
## 254 2019-09-11 0.01 21.314553 7.7500000 64.219840
## 255 2019-09-12 0.01 20.446271 7.7291667 40.623695
## 337 2019-12-03 0.01 216.156604 0.0100000 68.632945
## 338 2019-12-04 0.01 201.157492 0.0100000 7.432575
## 339 2019-12-05 0.01 6.402414 0.6666667 16.533794
## 428 2020-03-03 0.01 36.175165 7.5000000 73.256188
## 429 2020-03-04 0.01 30.714249 7.5000000 51.807804
## 441 2020-03-16 86608.30 46.242349 7.5000000 328.378090
## 443 2020-03-18 24116.79 162.072968 7.5000000 182.816099
## 470 2020-04-14 57405.42 196.065034 7.9791667 163.143392
## 471 2020-04-15 47372.83 61.180514 7.5000000 159.151775
## 543 2020-06-26 0.01 254.162244 8.0000000 165.895621
## 544 2020-06-27 0.01 251.012305 8.0000000 123.689598
## 545 2020-06-28 0.01 119.123441 8.0000000 172.755400
## 546 2020-06-29 0.01 213.881607 8.0000000 179.887378
## Corriente T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 ton_acum
## 15 0.0100000 40.653413 0.010000000 0.08846665 0.21740421 12797851.9
## 16 0.0100000 24.998914 0.010000000 0.01000000 0.01000000 12797851.9
## 17 0.0100000 22.536219 0.010000000 0.01000000 0.01000000 12797851.9
## 18 0.0100000 18.333527 0.010000000 0.01000000 0.01000000 12797851.9
## 19 0.0100000 19.000446 0.010000000 0.01000000 0.01000000 12797851.9
## 78 0.0619501 42.491642 0.012151670 1.54752564 1.13483974 17967475.8
## 79 0.0100000 38.936746 0.010000000 1.72880413 1.22959404 17967475.8
## 80 0.0100000 53.990010 0.010000000 4.30439486 4.51012512 17967475.8
## 81 0.9404026 51.224233 0.010000000 3.56563772 2.86478309 0.1
## 149 0.0100000 35.243278 0.011239038 1.57049968 1.10738804 5768427.0
## 176 0.0100000 30.350132 0.010000000 0.01000000 2.14505662 8019928.0
## 177 0.0100000 20.650585 0.010000000 0.01000000 0.10595621 8019683.0
## 178 0.0100000 10.699831 0.010000000 0.01000000 0.01000000 8019451.0
## 179 0.0100000 18.613152 0.010000000 0.01000000 0.01000000 8019190.0
## 180 0.0100000 22.922536 0.010000000 0.01000000 0.01000000 8018844.0
## 254 0.0100000 20.067897 0.010000000 0.01000000 0.01000000 14680160.0
## 255 0.5496111 52.085104 0.002948072 2.49409581 2.13951237 14679944.0
## 337 0.0100000 56.511999 0.010000000 3.69019072 2.71387467 21215100.0
## 338 0.0100000 57.343332 0.010000000 3.50762682 2.19957309 21214882.0
## 339 0.0100000 55.680846 0.010000000 3.54693676 3.25394602 0.1
## 428 0.0100000 18.562103 0.010000000 0.01000000 0.01000000 7487466.0
## 429 0.0100000 26.986597 0.010000000 0.27211529 0.09101781 7487291.0
## 441 125.3718041 57.601287 3.557042641 3.84038227 3.91625141 8524949.0
## 443 33.1940879 55.961543 0.730211426 3.33513297 3.41055237 8588329.0
## 470 53.0041524 57.569912 1.533047944 3.65087052 2.98557276 11180244.0
## 471 53.8091441 49.272328 1.692307869 3.26391592 2.49944402 11227616.0
## 543 0.0100000 6.889288 0.001132988 0.01000000 0.01000000 17639304.0
## 544 0.0100000 4.570233 0.010000000 0.01253734 0.01000000 17639010.0
## 545 0.0100000 4.330585 0.010000000 0.01000000 0.01000000 17638745.0
## 546 0.0100000 5.132844 0.010000000 0.01000000 0.01000000 17638464.0
head(Corriente_outliers, 30)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset
## 15 2019-01-15 0.01 225.93630 7.958333 135.151396
## 16 2019-01-16 0.01 202.84693 7.541667 8.556933
## 17 2019-01-17 0.01 202.03400 8.000000 8.611637
## 18 2019-01-18 0.01 201.25147 7.958333 8.742568
## 19 2019-01-19 0.01 53.47405 7.500000 157.181469
## 20 2019-01-20 36531.10 42.34228 7.500000 229.999434
## 43 2019-02-12 19958.07 112.94737 69.375000 222.760463
## 44 2019-02-13 19947.83 114.82612 7.500000 218.956877
## 77 2019-03-18 21195.62 198.49698 7.500000 219.192808
## 78 2019-03-19 0.01 22.05181 7.500000 34.429735
## 79 2019-03-20 0.01 0.01000 7.500000 0.010000
## 80 2019-03-21 0.01 0.01000 7.500000 0.010000
## 81 2019-03-22 0.01 18.59956 7.500000 31.421540
## 113 2019-04-23 42499.40 177.79559 7.500000 226.230017
## 114 2019-04-24 42498.04 177.48268 7.500000 225.185760
## 127 2019-05-07 27973.27 200.55711 7.500000 216.577735
## 149 2019-05-29 0.01 55.61887 7.500000 88.634211
## 150 2019-05-30 22511.35 53.99816 7.500000 209.187638
## 175 2019-06-24 19139.13 70.17309 7.500000 207.580697
## 176 2019-06-25 0.01 58.48542 7.500000 119.184292
## 177 2019-06-26 0.01 47.42849 7.500000 7.871639
## 178 2019-06-27 0.01 46.38774 7.500000 117.261253
## 179 2019-06-28 0.01 43.95892 7.500000 175.376969
## 180 2019-06-29 0.01 43.20602 7.500000 177.899915
## 226 2019-08-14 26770.06 207.58160 7.500000 219.750329
## 253 2019-09-10 30311.12 245.51034 7.750000 217.883641
## 254 2019-09-11 0.01 21.31455 7.750000 64.219840
## 255 2019-09-12 0.01 20.44627 7.729167 40.623695
## 281 2019-10-08 31308.55 125.87614 7.500000 228.451514
## 282 2019-10-09 20348.15 126.18405 7.500000 211.075542
## Corriente T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 ton_acum
## 15 0.0100000 40.65341 0.010000000 0.08846665 0.2174042 12797851.9
## 16 0.0100000 24.99891 0.010000000 0.01000000 0.0100000 12797851.9
## 17 0.0100000 22.53622 0.010000000 0.01000000 0.0100000 12797851.9
## 18 0.0100000 18.33353 0.010000000 0.01000000 0.0100000 12797851.9
## 19 0.0100000 19.00045 0.010000000 0.01000000 0.0100000 12797851.9
## 20 46.6191311 46.89799 5.659054995 1.70509991 2.2054545 12834383.0
## 43 51.7660758 37.33947 1.080412596 0.81720739 0.8285527 14978014.0
## 44 61.3042102 44.57258 0.735764346 1.90408600 1.5382229 14997961.8
## 77 53.7778369 52.24535 0.844764545 3.00683965 2.0315111 17967475.8
## 78 0.0619501 42.49164 0.012151670 1.54752564 1.1348397 17967475.8
## 79 0.0100000 38.93675 0.010000000 1.72880413 1.2295940 17967475.8
## 80 0.0100000 53.99001 0.010000000 4.30439486 4.5101251 17967475.8
## 81 0.9404026 51.22423 0.010000000 3.56563772 2.8647831 0.1
## 113 48.3859707 38.72933 1.665686992 1.24594335 1.4820104 2804110.0
## 114 61.2290119 33.56098 1.518362586 0.77955511 0.8241418 2846608.0
## 127 51.8322726 52.34146 1.219738799 2.21268675 1.9522210 4026882.0
## 149 0.0100000 35.24328 0.011239038 1.57049968 1.1073880 5768427.0
## 150 33.7159783 46.10239 0.682916357 2.98312126 2.1345206 5790938.0
## 175 43.5276566 35.97505 0.650703190 0.58664414 2.8296758 8020195.0
## 176 0.0100000 30.35013 0.010000000 0.01000000 2.1450566 8019928.0
## 177 0.0100000 20.65058 0.010000000 0.01000000 0.1059562 8019683.0
## 178 0.0100000 10.69983 0.010000000 0.01000000 0.0100000 8019451.0
## 179 0.0100000 18.61315 0.010000000 0.01000000 0.0100000 8019190.0
## 180 0.0100000 22.92254 0.010000000 0.01000000 0.0100000 8018844.0
## 226 42.3421670 42.14589 0.976805818 1.09626817 1.1339708 12267329.0
## 253 50.7894257 39.31355 0.991761042 1.18644027 1.3296140 14680440.0
## 254 0.0100000 20.06790 0.010000000 0.01000000 0.0100000 14680160.0
## 255 0.5496111 52.08510 0.002948072 2.49409581 2.1395124 14679944.0
## 281 49.6209374 38.83861 1.106780946 1.53540249 1.3535701 17039069.0
## 282 49.1702896 38.51890 0.513771874 0.82784962 0.9192092 17059417.0
head(T_aceite_motor_outliers, 30)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset
## 1 2019-01-01 154979.34 202.20612 7.500000 258.356143
## 2 2019-01-02 146932.65 204.22895 7.500000 251.452566
## 3 2019-01-03 127159.29 193.37900 7.500000 244.403748
## 15 2019-01-15 0.01 225.93630 7.958333 135.151396
## 16 2019-01-16 0.01 202.84693 7.541667 8.556933
## 17 2019-01-17 0.01 202.03400 8.000000 8.611637
## 18 2019-01-18 0.01 201.25147 7.958333 8.742568
## 19 2019-01-19 0.01 53.47405 7.500000 157.181469
## 43 2019-02-12 19958.07 112.94737 69.375000 222.760463
## 72 2019-03-13 91239.18 194.79082 7.500000 254.992007
## 79 2019-03-20 0.01 0.01000 7.500000 0.010000
## 113 2019-04-23 42499.40 177.79559 7.500000 226.230017
## 114 2019-04-24 42498.04 177.48268 7.500000 225.185760
## 148 2019-05-28 59331.13 133.27667 7.500000 226.481587
## 149 2019-05-29 0.01 55.61887 7.500000 88.634211
## 152 2019-06-01 134038.41 45.54174 7.500000 295.150701
## 153 2019-06-02 122063.63 42.44502 7.500000 297.394257
## 175 2019-06-24 19139.13 70.17309 7.500000 207.580697
## 176 2019-06-25 0.01 58.48542 7.500000 119.184292
## 177 2019-06-26 0.01 47.42849 7.500000 7.871639
## 178 2019-06-27 0.01 46.38774 7.500000 117.261253
## 179 2019-06-28 0.01 43.95892 7.500000 175.376969
## 180 2019-06-29 0.01 43.20602 7.500000 177.899915
## 197 2019-07-16 101684.93 105.25159 7.500000 293.098264
## 198 2019-07-17 87777.85 103.74011 7.500000 290.565069
## 199 2019-07-18 109022.51 110.10187 7.500000 293.782063
## 200 2019-07-19 116380.39 105.29876 7.500000 300.083221
## 201 2019-07-20 100037.73 109.64729 7.500000 293.916973
## 212 2019-07-31 99909.24 157.27203 7.500000 278.243835
## 213 2019-08-01 118022.97 154.01280 7.500000 311.095692
## Corriente T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 ton_acum
## 1 102.96132 37.69877 2.73399972 0.009995935 2.1021723 11437411
## 2 102.34145 36.84161 2.58082849 0.010000000 1.5989251 11584344
## 3 101.09568 38.76250 2.50820745 0.315446940 1.8541986 11711503
## 15 0.01000 40.65341 0.01000000 0.088466645 0.2174042 12797852
## 16 0.01000 24.99891 0.01000000 0.010000000 0.0100000 12797852
## 17 0.01000 22.53622 0.01000000 0.010000000 0.0100000 12797852
## 18 0.01000 18.33353 0.01000000 0.010000000 0.0100000 12797852
## 19 0.01000 19.00045 0.01000000 0.010000000 0.0100000 12797852
## 43 51.76608 37.33947 1.08041260 0.817207394 0.8285527 14978014
## 72 91.01126 31.28330 2.38401880 0.328016210 2.0868724 17531548
## 79 0.01000 38.93675 0.01000000 1.728804133 1.2295940 17967476
## 113 48.38597 38.72933 1.66568699 1.245943349 1.4820104 2804110
## 114 61.22901 33.56098 1.51836259 0.779555107 0.8241418 2846608
## 148 65.36430 29.24667 2.01541493 0.393102954 1.5123603 5768611
## 149 0.01000 35.24328 0.01123904 1.570499680 1.1073880 5768427
## 152 115.36232 23.62738 3.60508449 0.009529005 2.3450170 6016304
## 153 120.66392 36.01195 3.75285143 1.727273546 3.2016829 6138368
## 175 43.52766 35.97505 0.65070319 0.586644136 2.8296758 8020195
## 176 0.01000 30.35013 0.01000000 0.010000000 2.1450566 8019928
## 177 0.01000 20.65058 0.01000000 0.010000000 0.1059562 8019683
## 178 0.01000 10.69983 0.01000000 0.010000000 0.0100000 8019451
## 179 0.01000 18.61315 0.01000000 0.010000000 0.0100000 8019190
## 180 0.01000 22.92254 0.01000000 0.010000000 0.0100000 8018844
## 197 120.31462 40.18559 3.61791616 1.014019121 3.0554597 9557134
## 198 124.71366 24.30799 3.97676214 0.010000000 2.0977723 9644912
## 199 123.98846 25.84139 3.54888954 0.010000000 2.1322445 9753935
## 200 124.50532 14.22378 3.79331921 0.010000000 2.1950050 9870315
## 201 121.86375 40.50494 3.50229698 1.738521692 2.6532278 9970353
## 212 118.79870 39.75846 2.65432473 1.038834713 2.3295970 11115925
## 213 134.10914 40.46913 3.84922488 2.107395711 2.9892424 11233948
head(F80_CV_001_outliers, 30)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset
## 15 2019-01-15 0.010 225.93630 7.958333 135.151396
## 16 2019-01-16 0.010 202.84693 7.541667 8.556933
## 17 2019-01-17 0.010 202.03400 8.000000 8.611637
## 18 2019-01-18 0.010 201.25147 7.958333 8.742568
## 19 2019-01-19 0.010 53.47405 7.500000 157.181469
## 20 2019-01-20 36531.101 42.34228 7.500000 229.999434
## 44 2019-02-13 19947.831 114.82612 7.500000 218.956877
## 78 2019-03-19 0.010 22.05181 7.500000 34.429735
## 79 2019-03-20 0.010 0.01000 7.500000 0.010000
## 80 2019-03-21 0.010 0.01000 7.500000 0.010000
## 81 2019-03-22 0.010 18.59956 7.500000 31.421540
## 149 2019-05-29 0.010 55.61887 7.500000 88.634211
## 150 2019-05-30 22511.355 53.99816 7.500000 209.187638
## 175 2019-06-24 19139.135 70.17309 7.500000 207.580697
## 176 2019-06-25 0.010 58.48542 7.500000 119.184292
## 177 2019-06-26 0.010 47.42849 7.500000 7.871639
## 178 2019-06-27 0.010 46.38774 7.500000 117.261253
## 179 2019-06-28 0.010 43.95892 7.500000 175.376969
## 180 2019-06-29 0.010 43.20602 7.500000 177.899915
## 254 2019-09-11 0.010 21.31455 7.750000 64.219840
## 255 2019-09-12 0.010 20.44627 7.729167 40.623695
## 282 2019-10-09 20348.146 126.18405 7.500000 211.075542
## 295 2019-10-22 0.010 152.81046 7.500000 188.621991
## 296 2019-10-23 0.010 150.01135 7.500000 187.201082
## 297 2019-10-24 0.010 195.90201 7.500000 193.756401
## 298 2019-10-25 3496.097 228.90642 7.500000 201.193326
## 299 2019-10-26 0.010 203.56860 7.500000 199.499051
## 300 2019-10-27 0.010 203.05418 7.500000 201.385211
## 323 2019-11-19 17414.928 120.82477 7.500000 197.330050
## 336 2019-12-02 29696.147 206.97875 6.875000 220.744177
## Corriente T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 ton_acum
## 15 0.0100000 40.65341 0.010000000 0.08846665 0.2174042 12797851.9
## 16 0.0100000 24.99891 0.010000000 0.01000000 0.0100000 12797851.9
## 17 0.0100000 22.53622 0.010000000 0.01000000 0.0100000 12797851.9
## 18 0.0100000 18.33353 0.010000000 0.01000000 0.0100000 12797851.9
## 19 0.0100000 19.00045 0.010000000 0.01000000 0.0100000 12797851.9
## 20 46.6191311 46.89799 5.659054995 1.70509991 2.2054545 12834383.0
## 44 61.3042102 44.57258 0.735764346 1.90408600 1.5382229 14997961.8
## 78 0.0619501 42.49164 0.012151670 1.54752564 1.1348397 17967475.8
## 79 0.0100000 38.93675 0.010000000 1.72880413 1.2295940 17967475.8
## 80 0.0100000 53.99001 0.010000000 4.30439486 4.5101251 17967475.8
## 81 0.9404026 51.22423 0.010000000 3.56563772 2.8647831 0.1
## 149 0.0100000 35.24328 0.011239038 1.57049968 1.1073880 5768427.0
## 150 33.7159783 46.10239 0.682916357 2.98312126 2.1345206 5790938.0
## 175 43.5276566 35.97505 0.650703190 0.58664414 2.8296758 8020195.0
## 176 0.0100000 30.35013 0.010000000 0.01000000 2.1450566 8019928.0
## 177 0.0100000 20.65058 0.010000000 0.01000000 0.1059562 8019683.0
## 178 0.0100000 10.69983 0.010000000 0.01000000 0.0100000 8019451.0
## 179 0.0100000 18.61315 0.010000000 0.01000000 0.0100000 8019190.0
## 180 0.0100000 22.92254 0.010000000 0.01000000 0.0100000 8018844.0
## 254 0.0100000 20.06790 0.010000000 0.01000000 0.0100000 14680160.0
## 255 0.5496111 52.08510 0.002948072 2.49409581 2.1395124 14679944.0
## 282 49.1702896 38.51890 0.513771874 0.82784962 0.9192092 17059417.0
## 295 0.0100000 14.04247 0.010000000 0.01000000 0.0100000 18144009.0
## 296 0.0100000 13.44336 0.010000000 0.01000000 0.0100000 18143813.0
## 297 0.0100000 14.06387 0.010000000 0.01000000 0.0100000 18143616.0
## 298 7.3779928 24.77489 0.120462482 0.37936252 0.8047900 18147112.0
## 299 0.0100000 44.61288 0.010000000 0.95366430 1.3625010 18146941.0
## 300 0.4658632 49.83571 0.001651430 1.34516671 1.4137468 18146893.0
## 323 34.7202054 37.70118 0.541129081 0.64193697 0.7393594 20032795.0
## 336 44.7167507 55.88009 0.803888843 3.09165690 3.0095417 21215300.0
head(F80_CV_701_outliers, 30)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset Corriente
## 1 2019-01-01 154979.34 202.20612 7.500000 258.356143 102.9613
## 2 2019-01-02 146932.65 204.22895 7.500000 251.452566 102.3415
## 3 2019-01-03 127159.29 193.37900 7.500000 244.403748 101.0957
## 15 2019-01-15 0.01 225.93630 7.958333 135.151396 0.0100
## 16 2019-01-16 0.01 202.84693 7.541667 8.556933 0.0100
## 17 2019-01-17 0.01 202.03400 8.000000 8.611637 0.0100
## 18 2019-01-18 0.01 201.25147 7.958333 8.742568 0.0100
## 19 2019-01-19 0.01 53.47405 7.500000 157.181469 0.0100
## 152 2019-06-01 134038.41 45.54174 7.500000 295.150701 115.3623
## 176 2019-06-25 0.01 58.48542 7.500000 119.184292 0.0100
## 177 2019-06-26 0.01 47.42849 7.500000 7.871639 0.0100
## 178 2019-06-27 0.01 46.38774 7.500000 117.261253 0.0100
## 179 2019-06-28 0.01 43.95892 7.500000 175.376969 0.0100
## 180 2019-06-29 0.01 43.20602 7.500000 177.899915 0.0100
## 198 2019-07-17 87777.85 103.74011 7.500000 290.565069 124.7137
## 199 2019-07-18 109022.51 110.10187 7.500000 293.782063 123.9885
## 200 2019-07-19 116380.39 105.29876 7.500000 300.083221 124.5053
## 228 2019-08-16 87211.25 214.16643 7.500000 305.240606 135.6265
## 239 2019-08-27 116149.67 208.01346 7.750000 279.456718 118.1930
## 240 2019-08-28 92854.09 235.94889 7.750000 285.269579 124.8822
## 241 2019-08-29 128322.81 231.85832 7.750000 298.714832 127.1548
## 242 2019-08-30 127808.79 231.61351 7.750000 300.461638 128.9165
## 243 2019-08-31 117538.34 233.29323 7.750000 269.729842 110.0006
## 244 2019-09-01 123524.73 239.23400 7.750000 289.158970 128.1212
## 245 2019-09-02 116564.81 239.81485 7.750000 293.497420 120.9511
## 254 2019-09-11 0.01 21.31455 7.750000 64.219840 0.0100
## 270 2019-09-27 105041.02 91.89207 7.500000 315.130061 131.1363
## 271 2019-09-28 113204.00 92.20311 7.500000 309.759190 126.8006
## 295 2019-10-22 0.01 152.81046 7.500000 188.621991 0.0100
## 296 2019-10-23 0.01 150.01135 7.500000 187.201082 0.0100
## T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 ton_acum
## 1 37.698772 2.734000 0.009995935 2.1021723 11437411
## 2 36.841608 2.580828 0.010000000 1.5989251 11584344
## 3 38.762498 2.508207 0.315446940 1.8541986 11711503
## 15 40.653413 0.010000 0.088466645 0.2174042 12797852
## 16 24.998914 0.010000 0.010000000 0.0100000 12797852
## 17 22.536219 0.010000 0.010000000 0.0100000 12797852
## 18 18.333527 0.010000 0.010000000 0.0100000 12797852
## 19 19.000446 0.010000 0.010000000 0.0100000 12797852
## 152 23.627381 3.605084 0.009529005 2.3450170 6016304
## 176 30.350132 0.010000 0.010000000 2.1450566 8019928
## 177 20.650585 0.010000 0.010000000 0.1059562 8019683
## 178 10.699831 0.010000 0.010000000 0.0100000 8019451
## 179 18.613152 0.010000 0.010000000 0.0100000 8019190
## 180 22.922536 0.010000 0.010000000 0.0100000 8018844
## 198 24.307992 3.976762 0.010000000 2.0977723 9644912
## 199 25.841387 3.548890 0.010000000 2.1322445 9753935
## 200 14.223782 3.793319 0.010000000 2.1950050 9870315
## 228 52.532463 3.687398 5.080947604 3.4773882 12438888
## 239 30.479311 3.071697 0.001308198 1.9342927 13397776
## 240 15.599450 2.631546 0.010000000 1.6978125 13490630
## 241 10.598364 3.556292 0.010000000 2.5641800 13618953
## 242 9.585719 3.783953 0.010000000 2.6312161 13746762
## 243 10.404947 2.958800 0.010000000 2.0809340 13864300
## 244 21.897615 3.645474 0.010000000 1.8947872 13987825
## 245 23.437022 3.939584 0.010000000 2.2004913 14104390
## 254 20.067897 0.010000 0.010000000 0.0100000 14680160
## 270 24.855278 3.900000 0.050972159 0.7834499 16126896
## 271 11.069505 3.693499 0.010000000 1.8563040 16240100
## 295 14.042466 0.010000 0.010000000 0.0100000 18144009
## 296 13.443365 0.010000 0.010000000 0.0100000 18143813
head(F80_CV_005_outliers, 30)
## Timestamp T_H Posion_poste_manto CSS Presion_Hidroset
## 15 2019-01-15 0.010 225.93630 7.958333 135.151396
## 16 2019-01-16 0.010 202.84693 7.541667 8.556933
## 17 2019-01-17 0.010 202.03400 8.000000 8.611637
## 18 2019-01-18 0.010 201.25147 7.958333 8.742568
## 19 2019-01-19 0.010 53.47405 7.500000 157.181469
## 43 2019-02-12 19958.074 112.94737 69.375000 222.760463
## 114 2019-04-24 42498.036 177.48268 7.500000 225.185760
## 177 2019-06-26 0.010 47.42849 7.500000 7.871639
## 178 2019-06-27 0.010 46.38774 7.500000 117.261253
## 179 2019-06-28 0.010 43.95892 7.500000 175.376969
## 180 2019-06-29 0.010 43.20602 7.500000 177.899915
## 186 2019-07-05 102680.064 88.64622 7.500000 311.076490
## 221 2019-08-09 84050.517 194.46411 7.500000 275.036930
## 254 2019-09-11 0.010 21.31455 7.750000 64.219840
## 260 2019-09-17 124727.282 58.89216 7.500000 302.571574
## 270 2019-09-27 105041.021 91.89207 7.500000 315.130061
## 282 2019-10-09 20348.146 126.18405 7.500000 211.075542
## 294 2019-10-21 49180.597 158.24456 7.500000 236.615760
## 295 2019-10-22 0.010 152.81046 7.500000 188.621991
## 296 2019-10-23 0.010 150.01135 7.500000 187.201082
## 297 2019-10-24 0.010 195.90201 7.500000 193.756401
## 298 2019-10-25 3496.097 228.90642 7.500000 201.193326
## 323 2019-11-19 17414.928 120.82477 7.500000 197.330050
## 409 2020-02-13 1133.807 181.04041 7.500000 187.370872
## 428 2020-03-03 0.010 36.17516 7.500000 73.256188
## 429 2020-03-04 0.010 30.71425 7.500000 51.807804
## 506 2020-05-20 16923.401 227.27569 7.500000 201.207862
## 543 2020-06-26 0.010 254.16224 8.000000 165.895621
## 544 2020-06-27 0.010 251.01230 8.000000 123.689598
## 545 2020-06-28 0.010 119.12344 8.000000 172.755400
## Corriente T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005 ton_acum
## 15 0.010000 40.653413 0.010000000 0.08846665 0.21740421 12797852
## 16 0.010000 24.998914 0.010000000 0.01000000 0.01000000 12797852
## 17 0.010000 22.536219 0.010000000 0.01000000 0.01000000 12797852
## 18 0.010000 18.333527 0.010000000 0.01000000 0.01000000 12797852
## 19 0.010000 19.000446 0.010000000 0.01000000 0.01000000 12797852
## 43 51.766076 37.339466 1.080412596 0.81720739 0.82855274 14978014
## 114 61.229012 33.560984 1.518362586 0.77955511 0.82414185 2846608
## 177 0.010000 20.650585 0.010000000 0.01000000 0.10595621 8019683
## 178 0.010000 10.699831 0.010000000 0.01000000 0.01000000 8019451
## 179 0.010000 18.613152 0.010000000 0.01000000 0.01000000 8019190
## 180 0.010000 22.922536 0.010000000 0.01000000 0.01000000 8018844
## 186 127.468269 53.882645 3.610174377 3.61219249 0.54443975 8520782
## 221 118.254091 47.762328 1.059691715 1.23780387 5.99786403 11892992
## 254 0.010000 20.067897 0.010000000 0.01000000 0.01000000 14680160
## 260 121.669840 55.490344 4.772091615 4.13198716 5.04455086 15287750
## 270 131.136309 24.855278 3.899999572 0.05097216 0.78344989 16126896
## 282 49.170290 38.518901 0.513771874 0.82784962 0.91920920 17059417
## 294 61.597521 40.826797 1.657040397 1.95016406 0.99547379 18144197
## 295 0.010000 14.042466 0.010000000 0.01000000 0.01000000 18144009
## 296 0.010000 13.443365 0.010000000 0.01000000 0.01000000 18143813
## 297 0.010000 14.063865 0.010000000 0.01000000 0.01000000 18143616
## 298 7.377993 24.774885 0.120462482 0.37936252 0.80479005 18147112
## 323 34.720205 37.701183 0.541129081 0.64193697 0.73935943 20032795
## 409 2.691494 36.533235 0.059843660 0.71068078 0.64851458 5958095
## 428 0.010000 18.562103 0.010000000 0.01000000 0.01000000 7487466
## 429 0.010000 26.986597 0.010000000 0.27211529 0.09101781 7487291
## 506 24.074379 12.924443 0.554727372 0.01000000 0.15697862 14095368
## 543 0.010000 6.889288 0.001132988 0.01000000 0.01000000 17639304
## 544 0.010000 4.570233 0.010000000 0.01253734 0.01000000 17639010
## 545 0.010000 4.330585 0.010000000 0.01000000 0.01000000 17638745
head(ton_acum_outliers, 30)
## [1] Timestamp T_H Posion_poste_manto CSS
## [5] Presion_Hidroset Corriente T_aceite_motor F80_CV_001
## [9] F80_CV_701 F80_CV_005 ton_acum
## <0 rows> (o 0- extensión row.names)
T_H: No se consideran outliers ya que esta dentro de los parametros de posible procesamiento de mineral. Posion_poste_manto_outliers: No tiene outliers
CSS: Se imputara por modelos de regresion los valores que estan hacia arriba ya que hacia abajo el 0,01 coincide con dias del equipo detenido Se corrigen los valores ya que la apertura por motivos de geometría el CSS no podría ser mayor a 9in ó 10in. Consultado en FLSmidth.(2023). FLS TSUV gyratory crusher [Brochure]. https://www.flsmidth.com/-/media/brochures/brochures-products/crushing-and-sizing/2023/fls-tsuv-gyratory-crusher_brochure_es.pdf (Consultado el 20 de mayo de 2024). {Obtenido de las posibles relaciones entre CSS/OSS}
Presion_Hidroset: No se consideran outlier ya que puede estar dentro de los valores de trabajo Corriente: No se consideran outlier ya que puede estar dentro de los valores de trabajo T_aceite_motor: No se considera outlier por estar dentro de los posibles parametros F80_CV_001_outliers: F80_CV_701_outliers: En el F80 de las correas tranportadoras no se considerarán outlier por que los valores pueden representar valores puntuales en la operación F80_CV_005_outliers: ton_acum: No se consideran outliers
# Identificar los valores que son igual o mayor a 10
indices_a_imputar <- which(datosOpe1$CSS >= 10)
# Crear el modelo de regresión
# Usamos las demás columnas numéricas como variables independientes para predecir CSS
variables_independientes <- datosOpe1 %>%
select_if(is.numeric) %>%
select(-CSS)
# Separar los datos en entrenamiento y prueba
X <- variables_independientes[-indices_a_imputar, ]
y <- datosOpe1$CSS[-indices_a_imputar]
# Entrenar el modelo
modelo <- train(X, y, method = "lm")
# Realizar la imputación en los valores igual o mayor a 10
X_a_imputar <- variables_independientes[indices_a_imputar, ]
predicciones <- predict(modelo, X_a_imputar)
# Asignar las predicciones a la columna CSS
datosOpe1$CSS[indices_a_imputar] <- predicciones
# Mostrar los resultados
imputaciones <- datosOpe1[indices_a_imputar, c("Timestamp", "CSS")]
head(imputaciones)
## Timestamp CSS
## 42 2019-02-11 7.651395
## 43 2019-02-12 7.453195
## 70 2019-03-11 7.553947
## 71 2019-03-12 7.838107
## 92 2019-04-02 7.530160
## 93 2019-04-03 7.505214
Reducción de dimensionalidad
datosOpe2 <- datosOpe1
datosRepo2 <- datosRepo1
# Contar valores faltantes en datos de operacion
countNa7 <- sapply(datosOpe2, function(x) sum(is.na(x)))
kable(data.frame(Valores_Faltantes = countNa7), format = "markdown")
| Valores_Faltantes | |
|---|---|
| Timestamp | 0 |
| T_H | 0 |
| Posion_poste_manto | 0 |
| CSS | 0 |
| Presion_Hidroset | 0 |
| Corriente | 0 |
| T_aceite_motor | 0 |
| F80_CV_001 | 0 |
| F80_CV_701 | 0 |
| F80_CV_005 | 0 |
| ton_acum | 0 |
# Contar valores faltantes en datos de reportes
countNa8 <- sapply(datosRepo2, function(x) sum(is.na(x)))
kable(data.frame(Valores_Faltantes = countNa8), format = "markdown")
| Valores_Faltantes | |
|---|---|
| date_Time | 0 |
| Numcampana_med | 0 |
| diametro_Mto | 0 |
| dias_acum_Ccv | 0 |
| pos_med_Ccv | 0 |
| remanente_Ccv | 0 |
| desgaste_Ccv | 0 |
| tasa_desgaste_Ccv_acum | 0 |
| aplan_Ccv | 0 |
| pos_aplan_Ccv | 0 |
Existen 32 fechas únicas en los reportes de analisis dimensional, por lo tanto se considerarán ventanas temporales de 29 días, con la idea de cubrir las distintas caracterizaciones de los datos operacionales…
# Crear dataframes vacíos para almacenar los resultados
resultados_media_29 <- data.frame()
resultados_mediana_29 <- data.frame()
resultados_sd_29 <- data.frame()
# Iterar sobre las fechas únicas en datosRepo2
for(fecha in unique(datosRepo2$date_Time)) {
# Definir la ventana de 29 días alrededor de cada fecha única
inicio_ventana_29 <- fecha - 14
fin_ventana_29 <- fecha + 14
# Filtrar los datos de datosOpe2 dentro de la ventana de 29 días
datos_filtrados_29 <- datosOpe2 %>% filter(Timestamp >= inicio_ventana_29 & Timestamp <= fin_ventana_29)
# Excluir la columna Timestamp de los cálculos
datos_filtrados_29_sin_timestamp <- datos_filtrados_29 %>% select(-Timestamp)
# Calcular los estadísticos de los datos filtrados (29 días)
media_fila_29 <- datos_filtrados_29_sin_timestamp %>% summarise(across(everything(), ~mean(.x, na.rm = TRUE)))
mediana_fila_29 <- datos_filtrados_29_sin_timestamp %>% summarise(across(everything(), ~median(.x, na.rm = TRUE)))
sd_fila_29 <- datos_filtrados_29_sin_timestamp %>% summarise(across(everything(), ~sd(.x, na.rm = TRUE)))
# Agregar las fechas correspondientes (29 días)
media_fila_29$Fecha <- fecha
mediana_fila_29$Fecha <- fecha
sd_fila_29$Fecha <- fecha
# Agregar los resultados al dataframe de resultados (29 días)
resultados_media_29 <- rbind(resultados_media_29, media_fila_29)
resultados_mediana_29 <- rbind(resultados_mediana_29, mediana_fila_29)
resultados_sd_29 <- rbind(resultados_sd_29, sd_fila_29)
}
# Asignar los dataframes resultantes a variables
datosOpe2_media_29 <- resultados_media_29
datosOpe2_mediana_29 <- resultados_mediana_29
datosOpe2_sd_29 <- resultados_sd_29
# Convertir la columna 'Fecha' a formato de fecha en los datasets temporales
datosOpe2_media_29 <- datosOpe2_media_29 %>%
mutate(Fecha = as.Date(Fecha, origin = "1970-01-01"))
datosOpe2_mediana_29 <- datosOpe2_mediana_29 %>%
mutate(Fecha = as.Date(Fecha, origin = "1970-01-01"))
datosOpe2_sd_29 <- datosOpe2_sd_29 %>%
mutate(Fecha = as.Date(Fecha, origin = "1970-01-01"))
# Selección de columnas numéricas (excluyendo la columna de fecha y ton_acum)
cols <- names(datosOpe1)[sapply(datosOpe1, is.numeric) & names(datosOpe1) != "ton_acum"]
# Calcular estadísticas descriptivas utilizando todas las columnas numéricas
calcular_estadisticas_df <- function(data, cols, fun) {
data %>%
select(all_of(cols)) %>%
summarise(across(everything(), fun, na.rm = TRUE))
}
# Calcular estadísticas descriptivas utilizando la mediana
stats_original <- calcular_estadisticas_df(datosOpe1, cols, median)
## Warning: There was 1 warning in `summarise()`.
## ℹ In argument: `across(everything(), fun, na.rm = TRUE)`.
## Caused by warning:
## ! The `...` argument of `across()` is deprecated as of dplyr 1.1.0.
## Supply arguments directly to `.fns` through an anonymous function instead.
##
## # Previously
## across(a:b, mean, na.rm = TRUE)
##
## # Now
## across(a:b, \(x) mean(x, na.rm = TRUE))
stats_media <- calcular_estadisticas_df(datosOpe2_media_29, cols, median)
stats_mediana <- calcular_estadisticas_df(datosOpe2_mediana_29, cols, median)
stats_sd <- calcular_estadisticas_df(datosOpe2_sd_29, cols, median)
# Calcular IQR adicionalmente
iqr_original <- calcular_estadisticas_df(datosOpe1, cols, IQR)
iqr_media <- calcular_estadisticas_df(datosOpe2_media_29, cols, IQR)
iqr_mediana <- calcular_estadisticas_df(datosOpe2_mediana_29, cols, IQR)
iqr_sd <- calcular_estadisticas_df(datosOpe2_sd_29, cols, IQR)
# Unir estadísticas descriptivas
stats_combined <- bind_rows(
Original = stats_original,
Media_29 = stats_media,
Mediana_29 = stats_mediana,
SD_29 = stats_sd,
.id = "Dataset"
)
iqr_combined <- bind_rows(
Original = iqr_original,
Media_29 = iqr_media,
Mediana_29 = iqr_mediana,
SD_29 = iqr_sd,
.id = "Dataset"
)
print(stats_combined)
## Dataset T_H Posion_poste_manto CSS Presion_Hidroset Corriente
## 1 Original 90355.74 158.83998 7.5000000 255.58180 102.95397
## 2 Media_29 84894.66 144.69045 7.5112495 246.71742 94.64709
## 3 Mediana_29 90612.60 143.00871 7.5000000 253.01574 102.29015
## 4 SD_29 28977.78 37.28468 0.1249191 30.13322 27.24126
## T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005
## 1 53.443910 2.9690904 2.7128299 3.0321434
## 2 50.118491 2.7743564 2.5467268 2.9359382
## 3 53.814453 2.9972724 2.8385473 3.1137304
## 4 8.987474 0.9685264 0.9029723 0.8206935
print(iqr_combined)
## Dataset T_H Posion_poste_manto CSS Presion_Hidroset Corriente
## 1 Original 25808.357 109.95323 0.25000000 35.46671 19.49680
## 2 Media_29 12448.256 53.91476 0.16645474 25.16716 10.32892
## 3 Mediana_29 5877.126 72.54533 0.05729167 25.80002 10.65458
## 4 SD_29 13863.220 48.89781 0.21532383 51.08201 18.96842
## T_aceite_motor F80_CV_001 F80_CV_701 F80_CV_005
## 1 5.445500 1.0282536 1.1552856 1.0042504
## 2 4.704325 0.4206991 0.4226124 0.5605405
## 3 2.227679 0.3655583 0.4450567 0.6049711
## 4 7.025113 0.4177410 0.3931106 0.3815825
# Crear datos combinados para gráficos
datos_combined <- bind_rows(
datosOpe1 %>% rename(Fecha = Timestamp) %>% mutate(Dataset = "Original"),
datosOpe2_media_29 %>% mutate(Dataset = "Media_29"),
datosOpe2_mediana_29 %>% mutate(Dataset = "Mediana_29"),
datosOpe2_sd_29 %>% mutate(Dataset = "SD_29")
)
# Función para graficar líneas de tiempo
plot_line <- function(data, column, title) {
ggplot(data, aes(x = Fecha, y = !!sym(column), color = Dataset, alpha = Dataset, size = Dataset)) +
geom_line() +
scale_alpha_manual(values = c("Original" = 0.6, "Media_29" = 1, "Mediana_29" = 1, "SD_29" = 1)) +
scale_size_manual(values = c("Original" = 0.5, "Media_29" = 0.7, "Mediana_29" = 0.7, "SD_29" = 0.7)) +
labs(title = title, x = "Fecha", y = column) +
theme_minimal()
}
# Función para graficar boxplots
plot_boxplot <- function(data, column, title) {
ggplot(data, aes(x = Dataset, y = !!sym(column), fill = Dataset)) +
geom_boxplot() +
labs(title = title, x = "Dataset", y = column) +
theme_minimal()
}
# Función para graficar histogramas
plot_histogram <- function(data, column, title) {
ggplot(data, aes(x = !!sym(column), fill = Dataset)) +
geom_histogram(alpha = 0.6, position = "dodge", bins = 30) +
labs(title = title, x = column, y = "Count") +
theme_minimal()
}
# Mostrar gráficos para cada columna
for (column in cols) {
print(plot_line(datos_combined, column, paste(column, "over Time")))
print(plot_boxplot(datos_combined, column, paste("Boxplot of", column)))
print(plot_histogram(datos_combined, column, paste("Histogram of", column)))
}
# Función para calcular las diferencias en las estadísticas descriptivas
diff_stats <- function(data1, data2, cols) {
diffs <- lapply(cols, function(column) {
diff_mean <- abs(mean(data1[[column]], na.rm = TRUE) - mean(data2[[column]], na.rm = TRUE))
diff_median <- abs(median(data1[[column]], na.rm = TRUE) - median(data2[[column]], na.rm = TRUE))
diff_sd <- abs(sd(data1[[column]], na.rm = TRUE) - sd(data2[[column]], na.rm = TRUE))
diff_iqr <- abs(IQR(data1[[column]], na.rm = TRUE) - IQR(data2[[column]], na.rm = TRUE))
c(mean = diff_mean, median = diff_median, sd = diff_sd, iqr = diff_iqr)
})
diffs_df <- as.data.frame(do.call(rbind, diffs))
colnames(diffs_df) <- c("mean_diff", "median_diff", "sd_diff", "iqr_diff")
rownames(diffs_df) <- cols
return(diffs_df)
}
# Calcular diferencias para cada comparativa
diff_media <- diff_stats(datosOpe1, datosOpe2_media_29, cols)
diff_mediana <- diff_stats(datosOpe1, datosOpe2_mediana_29, cols)
diff_sd <- diff_stats(datosOpe1, datosOpe2_sd_29, cols)
# Crear un tibble para almacenar las diferencias de estadísticas descriptivas
diff_combined_media <- as.data.frame(diff_media)
diff_combined_mediana <- as.data.frame(diff_mediana)
diff_combined_sd <- as.data.frame(diff_sd)
# Redondear los valores para mejor legibilidad
diff_combined_media <- diff_combined_media %>% mutate(across(everything(), round, 2))
diff_combined_mediana <- diff_combined_mediana %>% mutate(across(everything(), round, 2))
diff_combined_sd <- diff_combined_sd %>% mutate(across(everything(), round, 2))
# Calcular la suma de las diferencias absolutas para cada fila individualmente
calculate_total_diff <- function(data) {
data %>%
mutate(total_diff = rowSums(across(mean_diff:iqr_diff)))
}
# Calcular total_diff y agregarlo a cada tabla de diferencias
diff_combined_media <- calculate_total_diff(diff_combined_media)
diff_combined_mediana <- calculate_total_diff(diff_combined_mediana)
diff_combined_sd <- calculate_total_diff(diff_combined_sd)
# Crear una fila de total para cada dataset
add_total_row <- function(data, dataset_name) {
total_diff <- sum(data$total_diff, na.rm = TRUE)
total_row <- data.frame(mean_diff = NA, median_diff = NA, sd_diff = NA, iqr_diff = NA, total_diff = total_diff, row.names = paste0("Total_", dataset_name))
bind_rows(data, total_row)
}
# Agregar fila de total a cada tabla de diferencias
diff_combined_media <- add_total_row(diff_combined_media, "Media_29")
diff_combined_mediana <- add_total_row(diff_combined_mediana, "Mediana_29")
diff_combined_sd <- add_total_row(diff_combined_sd, "SD_29")
# Mostrar diferencias absolutas para Media_29
diff_combined_media %>%
kable("html", caption = "Diferencias Absolutas para Media_29") %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
column_spec(1, bold = TRUE, color = "white", background = "dodgerblue")
| mean_diff | median_diff | sd_diff | iqr_diff | total_diff | |
|---|---|---|---|---|---|
| T_H | 990.53 | 5461.08 | 20617.76 | 13360.10 | 40429.47 |
| Posion_poste_manto | 7.06 | 14.15 | 26.17 | 56.04 | 103.42 |
| CSS | 0.05 | 0.01 | 0.41 | 0.08 | 0.55 |
| Presion_Hidroset | 3.18 | 8.86 | 27.91 | 10.30 | 50.25 |
| Corriente | 1.68 | 8.31 | 20.01 | 9.17 | 39.17 |
| T_aceite_motor | 0.03 | 3.33 | 7.14 | 0.74 | 11.24 |
| F80_CV_001 | 0.02 | 0.19 | 0.66 | 0.61 | 1.48 |
| F80_CV_701 | 0.04 | 0.17 | 0.67 | 0.73 | 1.61 |
| F80_CV_005 | 0.01 | 0.10 | 0.52 | 0.44 | 1.07 |
| Total_Media_29 | NA | NA | NA | NA | 40638.26 |
# Mostrar diferencias absolutas para Mediana_29
diff_combined_mediana %>%
kable("html", caption = "Diferencias Absolutas para Mediana_29") %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
column_spec(1, bold = TRUE, color = "white", background = "dodgerblue")
| mean_diff | median_diff | sd_diff | iqr_diff | total_diff | |
|---|---|---|---|---|---|
| T_H | 5484.48 | 256.86 | 22358.07 | 19931.23 | 48030.64 |
| Posion_poste_manto | 12.74 | 15.83 | 14.31 | 37.41 | 80.29 |
| CSS | 0.05 | 0.00 | 0.57 | 0.19 | 0.81 |
| Presion_Hidroset | 6.87 | 2.57 | 31.51 | 9.67 | 50.62 |
| Corriente | 6.27 | 0.66 | 20.75 | 8.84 | 36.52 |
| T_aceite_motor | 3.17 | 0.37 | 8.03 | 3.22 | 14.79 |
| F80_CV_001 | 0.21 | 0.03 | 0.68 | 0.66 | 1.58 |
| F80_CV_701 | 0.27 | 0.13 | 0.66 | 0.71 | 1.77 |
| F80_CV_005 | 0.14 | 0.08 | 0.54 | 0.40 | 1.16 |
| Total_Mediana_29 | NA | NA | NA | NA | 48218.18 |
# Mostrar diferencias absolutas para SD_29
diff_combined_sd %>%
kable("html", caption = "Diferencias Absolutas para SD_29") %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
column_spec(1, bold = TRUE, color = "white", background = "dodgerblue")
| mean_diff | median_diff | sd_diff | iqr_diff | total_diff | |
|---|---|---|---|---|---|
| T_H | 55237.49 | 61377.96 | 21400.16 | 11945.14 | 149960.75 |
| Posion_poste_manto | 114.80 | 121.56 | 35.74 | 61.06 | 333.16 |
| CSS | 7.10 | 7.38 | 0.07 | 0.03 | 14.58 |
| Presion_Hidroset | 207.17 | 225.45 | 18.85 | 15.62 | 467.09 |
| Corriente | 69.12 | 75.71 | 18.21 | 0.53 | 163.57 |
| T_aceite_motor | 41.20 | 44.46 | 5.71 | 1.58 | 92.95 |
| F80_CV_001 | 1.80 | 2.00 | 0.75 | 0.61 | 5.16 |
| F80_CV_701 | 1.57 | 1.81 | 0.77 | 0.76 | 4.91 |
| F80_CV_005 | 2.09 | 2.21 | 0.69 | 0.62 | 5.61 |
| Total_SD_29 | NA | NA | NA | NA | 151047.78 |
# Crear un resumen total
total_diff_summary <- bind_rows(
diff_combined_media %>% filter(row.names(diff_combined_media) == "Total_Media_29") %>% select(total_diff) %>% mutate(Dataset = "Media_29"),
diff_combined_mediana %>% filter(row.names(diff_combined_mediana) == "Total_Mediana_29") %>% select(total_diff) %>% mutate(Dataset = "Mediana_29"),
diff_combined_sd %>% filter(row.names(diff_combined_sd) == "Total_SD_29") %>% select(total_diff) %>% mutate(Dataset = "SD_29")
)
# Redondear los valores del resumen total
total_diff_summary <- total_diff_summary %>% mutate(total_diff = round(total_diff, 2))
# Mostrar resumen total de diferencias absolutas
total_diff_summary %>%
kable("html", caption = "Resumen de Diferencias Absolutas Totales") %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
column_spec(1, bold = TRUE, color = "white", background = "dodgerblue") %>%
column_spec(2, bold = TRUE, color = "white", background = "dodgerblue")
| total_diff | Dataset | |
|---|---|---|
| Total_Media_29 | 40638.26 | Media_29 |
| Total_Mediana_29 | 48218.18 | Mediana_29 |
| Total_SD_29 | 151047.78 | SD_29 |
# Hacer el match por fecha y agregar la columna ton_acum2
datosOpe2_mediana_29 <- merge(datosOpe2_mediana_29, datosOpe1[, c("Timestamp", "ton_acum")], by.x = "Fecha", by.y = "Timestamp", all.x = TRUE)
# Renombrar las columnas correctamente
colnames(datosOpe2_mediana_29)[colnames(datosOpe2_mediana_29) == "ton_acum.x"] <- "ton_acum"
colnames(datosOpe2_mediana_29)[colnames(datosOpe2_mediana_29) == "ton_acum.y"] <- "ton_acum2"
# Utilizando solo las observaciones con las medidas 6,5,4,3
datosRepo2 <- datosRepo2[datosRepo2$pos_med_Ccv %in% c(6, 5, 4, 3), ]
datosRepo2 <- datosRepo2 %>%
group_by(date_Time) %>%
summarise(
diametro_Mto = first(diametro_Mto),
dias_acum_Ccv = first(dias_acum_Ccv),
remanente_Ccv = mean(remanente_Ccv, na.rm = TRUE),
desgaste_Ccv = mean(desgaste_Ccv, na.rm = TRUE),
tasa_desgaste_Ccv_acum = mean(tasa_desgaste_Ccv_acum, na.rm = TRUE),
aplan_Ccv = mean(aplan_Ccv, na.rm = TRUE)
)
# Ver el resultado
head(datosRepo2,30)
## # A tibble: 30 × 7
## date_Time diametro_Mto dias_acum_Ccv remanente_Ccv desgaste_Ccv
## <date> <chr> <dbl> <dbl> <dbl>
## 1 2019-01-20 115 149 81.8 58.2
## 2 2019-02-13 115 173 74.6 65.3
## 3 2019-03-18 115 206 63.6 76.3
## 4 2019-03-22 110 0 140. 0
## 5 2019-04-23 110 32 126. 13.6
## 6 2019-06-04 113 74 113. 26.7
## 7 2019-06-24 113 94 101. 38.8
## 8 2019-08-06 113 137 89.0 51.0
## 9 2019-09-12 115 174 77.8 62.2
## 10 2019-10-09 115 201 68.6 71.4
## # ℹ 20 more rows
## # ℹ 2 more variables: tasa_desgaste_Ccv_acum <dbl>, aplan_Ccv <dbl>
columns_to_replace <- c("dias_acum_Ccv", "tasa_desgaste_Ccv_acum", "desgaste_Ccv")
datosRepo2[columns_to_replace] <- lapply(datosRepo2[columns_to_replace], function(x) {
x[x == 0] <- 0.01
return(x)
})
# Calcular el valor de desplazamiento
shift_value <- abs(min(datosRepo2$aplan_Ccv, na.rm = TRUE)) + 1
# Aplicar el desplazamiento a toda la columna 'aplan_Ccv'
datosRepo2 <- datosRepo2 %>%
mutate(Shift_aplan_Ccv = aplan_Ccv + shift_value)
pos_med_Ccv - pos_aplan_Ccv: Como se obtuvieron reducciones por ventanas temporales estas variables ya no aportan la informacion objetiva por lo que serán eliminadas.
Numcampana_med: No aporta informacion adicional
remanente_Ccv: Se evaluara mediante un análisis la opción de
eliminar, este valor se obtiene de la siguiente formula remanente_Ccv =
Espesor Nominal - desgaste_Ccv
Si la correlación es muy alta se decidirá imputar por que incluirla en
el modelo podría ocultar o distorsionar el impacto de otras variables
operacionales que también afectan el desgaste.
dias_acum_Ccv : Representa los dias de avance de la campaña, los días 0 se reemplazarán por 0,01 desgaste_Ccv : Lo desgaste que representan 0, indican que la condicion de cóncavas eran nuevas, para efecto de mejoras de análisis se reemplazará el 0 por 0,01 tasa_desgaste_Ccv_acum : Este valor representa la division de el desgaste (mm) en el tonelaje acumulado por MM de toneladas. El 0 se reemplazará po 0,01
Shift a columna aplan_Ccv
Se “desplaza” el valor, primero se identifica el valor minimo de la columna, se aplica valor absoluto y se suma el valor 1 obteniendo de esta manera nuestro Shift Value
Shift Value = | Mínimo valor | + 1
5,9 = | -4,9 | + 1
La nueva columna se llamará Shift_aplan_Ccv
Se crearan las siguientes variables a partir de las siguientes consideraciones.
ratio_presion_corriente: Presion_hidroset / Corriente ratio_carga_corriente: T_H / Corriente ratio_mto_css: posion_poste_manto / CSS ratio_th_mto: T_H / Posion_poste_manto ratio_th_cp: T_H / Corriente + Presion_Hidroset
# Asegurándonos de trabajar con las variables numéricas apropiadas
datos_numericos <- datosRepo2[, c("dias_acum_Ccv", "remanente_Ccv", "desgaste_Ccv", "tasa_desgaste_Ccv_acum", "aplan_Ccv", "Shift_aplan_Ccv")]
# Calcular la matriz de correlación
correlaciones <- cor(datos_numericos, use = "complete.obs")
print(correlaciones)
## dias_acum_Ccv remanente_Ccv desgaste_Ccv
## dias_acum_Ccv 1.0000000 -0.9929678 0.9930557
## remanente_Ccv -0.9929678 1.0000000 -0.9999952
## desgaste_Ccv 0.9930557 -0.9999952 1.0000000
## tasa_desgaste_Ccv_acum 0.5279688 -0.5585276 0.5583038
## aplan_Ccv -0.5648395 0.5947661 -0.5936715
## Shift_aplan_Ccv -0.5648395 0.5947661 -0.5936715
## tasa_desgaste_Ccv_acum aplan_Ccv Shift_aplan_Ccv
## dias_acum_Ccv 0.52796885 -0.56483946 -0.56483946
## remanente_Ccv -0.55852759 0.59476612 0.59476612
## desgaste_Ccv 0.55830384 -0.59367155 -0.59367155
## tasa_desgaste_Ccv_acum 1.00000000 -0.09070859 -0.09070859
## aplan_Ccv -0.09070859 1.00000000 1.00000000
## Shift_aplan_Ccv -0.09070859 1.00000000 1.00000000
# Transformar la matriz de correlación para el graficado
correlaciones_melt <- melt(correlaciones)
# Crear el mapa de calor
ggplot(data = correlaciones_melt, aes(x = Var1, y = Var2, fill = value)) +
geom_tile() +
scale_fill_gradient2(low = "blue", high = "red", mid = "white", midpoint = 0, limit = c(-1,1), space = "Lab", name="Correlación") +
theme_minimal() +
theme(axis.text.x = element_text(angle = 45, vjust = 1, size = 12, hjust = 1),
axis.text.y = element_text(size = 12)) +
geom_text(aes(label = sprintf("%.2f", value)), color = "black", size = 4) +
labs(x = "Variables", y = "Variables", title = "Mapa de calor de correlación")
# Eliminar columnas se justifica "remanente_Ccv" mediante la alta correlacion= 1.0
datosRepo2 <- datosRepo2[, !names(datosRepo2) %in% c("pos_med_Ccv", "pos_aplan_Ccv", "Numcampana_med", "aplan_Ccv", "remanente_Ccv", "dias_acum_Ccv")]
### Ingenieria de caracteristicas (creación de nuevas variables a partir de las existentes)
datosOpe3 <- datosOpe2_mediana_29
datosRepo3 <- datosRepo2
# Creando nuevas características datosOpe3
datosOpe3$ratio_presion_corriente <- datosOpe3$Presion_Hidroset / datosOpe3$Corriente
datosOpe3$ratio_carga_corriente <- datosOpe3$T_H / datosOpe3$Corriente
datosOpe3$ratio_mto_css <- datosOpe3$Posion_poste_manto / datosOpe3$CSS
datosOpe3$ratio_th_mto <- datosOpe3$T_H / datosOpe3$Posion_poste_manto
datosOpe3$ratio_th_cp <- datosOpe3$T_H / (datosOpe3$Corriente + datosOpe3$Presion_Hidroset)
datosOpe4 <- datosOpe3
datosRepo4 <- datosRepo3
# Función para aplicar y guardar transformaciones logarítmicas directamente en el dataframe original
apply_and_save_log_transform <- function(dataframe) {
# Evaluar asimetría y curtosis y aplicar logaritmo si es necesario
log_transformed_dataframe <- dataframe %>%
mutate(across(where(is.numeric), ~ {
skewness_value <- skewness(., na.rm = TRUE)
kurtosis_value <- kurtosis(., na.rm = TRUE)
if (skewness_value > 1 || kurtosis_value > 3) {
ifelse(. <= 0, NA, log(.))
} else {
.
}
}, .names = "log_{.col}"))
# Identificar columnas que fueron transformadas
transformed_columns <- names(select(log_transformed_dataframe, starts_with("log_")))
original_columns <- gsub("log_", "", transformed_columns)
# Eliminar columnas originales que fueron transformadas
log_transformed_dataframe <- log_transformed_dataframe %>%
select(-all_of(original_columns))
# Evaluar la necesidad de binning
binning_info <- sapply(select_if(log_transformed_dataframe, is.numeric), function(x) {
unique_values <- length(unique(na.omit(x)))
if (unique_values < 10) {
return(paste("Binning recomendado:", unique_values, "valores únicos después de la transformación logarítmica."))
} else {
return("No se recomienda binning.")
}
})
list(dataframe = log_transformed_dataframe, binning = binning_info)
}
# Aplicar la función
result_datosOpe4 <- apply_and_save_log_transform(datosOpe4)
# Actualizar
datosOpe4 <- result_datosOpe4$dataframe
datosOpe5 <- datosOpe4
datosRepo5 <- datosRepo4
View(datosOpe5)
View(datosRepo5)
# Función para generar y mostrar los boxplots en grillas de 2x3
generar_boxplots <- function(datos, titulo_dataset) {
datos_numericos <- datos[, sapply(datos, is.numeric)]
plots <- lapply(colnames(datos_numericos), function(col) {
ggplot(datos, aes_string(y = col)) +
geom_boxplot() +
ggtitle(paste("Boxplot de", col, "(", titulo_dataset, ")")) +
theme(axis.text.x = element_blank(), axis.ticks.x = element_blank())
})
num_plots <- length(plots)
plots_per_page <- 6
num_pages <- ceiling(num_plots / plots_per_page)
for (i in seq_len(num_pages)) {
start_index <- (i - 1) * plots_per_page + 1
end_index <- min(i * plots_per_page, num_plots)
do.call("grid.arrange", c(plots[start_index:end_index], ncol = 2, nrow = 3))
}
}
# Generar boxplots para datosOpe5 y datosRepo5
generar_boxplots(datosOpe5, "datosOpe5")
## Warning: `aes_string()` was deprecated in ggplot2 3.0.0.
## ℹ Please use tidy evaluation idioms with `aes()`.
## ℹ See also `vignette("ggplot2-in-packages")` for more information.
## This warning is displayed once every 8 hours.
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was
## generated.
generar_boxplots(datosRepo5, "datosRepo5")
# Función para generar y mostrar gráficos de líneas de tiempo en grillas de 2x3
generar_lineas_tiempo <- function(datos, fecha_col, titulo_dataset) {
datos_numericos <- datos[, sapply(datos, is.numeric)]
plots <- lapply(colnames(datos_numericos), function(col) {
ggplot(datos, aes_string(x = fecha_col, y = col)) +
geom_line() +
ggtitle(paste("Lineas de tiempo de", col, "(", titulo_dataset, ")")) +
theme(axis.text.x = element_text(angle = 90, hjust = 1))
})
num_plots <- length(plots)
plots_per_page <- 6
num_pages <- ceiling(num_plots / plots_per_page)
for (i in seq_len(num_pages)) {
start_index <- (i - 1) * plots_per_page + 1
end_index <- min(i * plots_per_page, num_plots)
do.call("grid.arrange", c(plots[start_index:end_index], ncol = 2, nrow = 3))
}
}
# Generar y mostrar las líneas de tiempo para datosOpe5 y datosRepo5
generar_lineas_tiempo(datosOpe5, "Fecha", "datosOpe5")
generar_lineas_tiempo(datosRepo5, "date_Time", "datosRepo5")
# Función para crear y mostrar heatmaps de correlación
crear_heatmap_correlacion <- function(datos, titulo) {
datos_numericos <- datos[sapply(datos, is.numeric)]
correlation_matrix <- cor(datos_numericos, use = "complete.obs")
melted_correlation <- melt(correlation_matrix)
ggplot(data = melted_correlation, aes(x = Var1, y = Var2, fill = value)) +
geom_tile() +
scale_fill_gradient2(low = "blue", high = "red", mid = "white", midpoint = 0,
limit = c(-1, 1), space = "Lab", name = "Correlación") +
theme_minimal() +
theme(axis.text.x = element_text(angle = 45, vjust = 1, size = 12, hjust = 1)) +
coord_fixed() +
ggtitle(titulo)
}
# Crear y mostrar los heatmaps para datosOpe5 y datosRepo5
heatmap_ope <- crear_heatmap_correlacion(datosOpe5, "Heatmap de Correlación - datosOpe5")
heatmap_repo <- crear_heatmap_correlacion(datosRepo5, "Heatmap de Correlación - datosRepo5")
print(heatmap_ope)
print(heatmap_repo)
# Función para calcular estadísticas
calcular_estadisticas <- function(datos) {
datos %>%
summarise(across(where(is.numeric), list(
media = ~mean(.x, na.rm = TRUE),
mediana = ~median(.x, na.rm = TRUE),
desviacion = ~sd(.x, na.rm = TRUE),
min = ~min(.x, na.rm = TRUE),
max = ~max(.x, na.rm = TRUE),
IQR = ~IQR(.x, na.rm = TRUE)
), .names = "{col}_{fn}"))
}
# Formatear los valores para mayor legibilidad
formatear_estadisticas <- function(df) {
df %>%
mutate(across(everything(), ~format(round(.x, 2), nsmall = 2, scientific = FALSE)))
}
# Calcular y formatear estadísticas básicas para datosOpe5 y datosRepo5
estadisticas_ope <- calcular_estadisticas(datosOpe5) %>% formatear_estadisticas()
estadisticas_repo <- calcular_estadisticas(datosRepo5) %>% formatear_estadisticas()
# Transformar y mostrar las estadísticas básicas en tablas con kable
mostrar_estadisticas <- function(estadisticas, titulo) {
estadisticas %>%
t() %>%
as.data.frame() %>%
rownames_to_column("Estadistica") %>%
kable(col.names = c("Estadística", "Valor"), caption = titulo)
}
# Mostrar tablas de estadísticas
mostrar_estadisticas(estadisticas_ope, "Estadísticas - datosOpe5")
| Estadística | Valor |
|---|---|
| log_T_H_media | 11.40 |
| log_T_H_mediana | 11.41 |
| log_T_H_desviacion | 0.08 |
| log_T_H_min | 11.09 |
| log_T_H_max | 11.49 |
| log_T_H_IQR | 0.07 |
| log_Posion_poste_manto_media | 148.09 |
| log_Posion_poste_manto_mediana | 143.01 |
| log_Posion_poste_manto_desviacion | 50.54 |
| log_Posion_poste_manto_min | 70.92 |
| log_Posion_poste_manto_max | 257.79 |
| log_Posion_poste_manto_IQR | 72.55 |
| log_CSS_media | 2.03 |
| log_CSS_mediana | 2.01 |
| log_CSS_desviacion | 0.02 |
| log_CSS_min | 2.01 |
| log_CSS_max | 2.08 |
| log_CSS_IQR | 0.01 |
| log_Presion_Hidroset_media | 256.83 |
| log_Presion_Hidroset_mediana | 253.02 |
| log_Presion_Hidroset_desviacion | 14.51 |
| log_Presion_Hidroset_min | 234.89 |
| log_Presion_Hidroset_max | 287.35 |
| log_Presion_Hidroset_IQR | 25.80 |
| log_Corriente_media | 102.85 |
| log_Corriente_mediana | 102.29 |
| log_Corriente_desviacion | 8.21 |
| log_Corriente_min | 88.62 |
| log_Corriente_max | 123.83 |
| log_Corriente_IQR | 10.65 |
| log_T_aceite_motor_media | 3.97 |
| log_T_aceite_motor_mediana | 3.99 |
| log_T_aceite_motor_desviacion | 0.04 |
| log_T_aceite_motor_min | 3.87 |
| log_T_aceite_motor_max | 4.04 |
| log_T_aceite_motor_IQR | 0.04 |
| log_F80_CV_001_media | 1.09 |
| log_F80_CV_001_mediana | 1.10 |
| log_F80_CV_001_desviacion | 0.10 |
| log_F80_CV_001_min | 0.84 |
| log_F80_CV_001_max | 1.30 |
| log_F80_CV_001_IQR | 0.12 |
| log_F80_CV_701_media | 1.02 |
| log_F80_CV_701_mediana | 1.04 |
| log_F80_CV_701_desviacion | 0.13 |
| log_F80_CV_701_min | 0.68 |
| log_F80_CV_701_max | 1.28 |
| log_F80_CV_701_IQR | 0.16 |
| log_F80_CV_005_media | 3.08 |
| log_F80_CV_005_mediana | 3.11 |
| log_F80_CV_005_desviacion | 0.39 |
| log_F80_CV_005_min | 2.31 |
| log_F80_CV_005_max | 3.79 |
| log_F80_CV_005_IQR | 0.60 |
| log_ton_acum_media | 10453470.41 |
| log_ton_acum_mediana | 10815622.50 |
| log_ton_acum_desviacion | 6596044.38 |
| log_ton_acum_min | 977769.00 |
| log_ton_acum_max | 22232700.00 |
| log_ton_acum_IQR | 10910838.42 |
| log_ton_acum2_media | 10398671.75 |
| log_ton_acum2_mediana | 10815740.00 |
| log_ton_acum2_desviacion | 6962644.89 |
| log_ton_acum2_min | 0.10 |
| log_ton_acum2_max | 23172738.00 |
| log_ton_acum2_IQR | 10918523.60 |
| log_ratio_presion_corriente_media | 0.92 |
| log_ratio_presion_corriente_mediana | 0.91 |
| log_ratio_presion_corriente_desviacion | 0.04 |
| log_ratio_presion_corriente_min | 0.84 |
| log_ratio_presion_corriente_max | 1.00 |
| log_ratio_presion_corriente_IQR | 0.03 |
| log_ratio_carga_corriente_media | 6.77 |
| log_ratio_carga_corriente_mediana | 6.79 |
| log_ratio_carga_corriente_desviacion | 0.11 |
| log_ratio_carga_corriente_min | 6.46 |
| log_ratio_carga_corriente_max | 6.93 |
| log_ratio_carga_corriente_IQR | 0.13 |
| log_ratio_mto_css_media | 19.44 |
| log_ratio_mto_css_mediana | 19.07 |
| log_ratio_mto_css_desviacion | 6.32 |
| log_ratio_mto_css_min | 9.46 |
| log_ratio_mto_css_max | 32.22 |
| log_ratio_mto_css_IQR | 9.47 |
| log_ratio_th_mto_media | 677.81 |
| log_ratio_th_mto_mediana | 676.19 |
| log_ratio_th_mto_desviacion | 239.51 |
| log_ratio_th_mto_min | 264.10 |
| log_ratio_th_mto_max | 1298.55 |
| log_ratio_th_mto_IQR | 366.21 |
| log_ratio_th_cp_media | 5.51 |
| log_ratio_th_cp_mediana | 5.53 |
| log_ratio_th_cp_desviacion | 0.10 |
| log_ratio_th_cp_min | 5.25 |
| log_ratio_th_cp_max | 5.69 |
| log_ratio_th_cp_IQR | 0.11 |
mostrar_estadisticas(estadisticas_repo, "Estadísticas - datosRepo5")
| Estadística | Valor |
|---|---|
| desgaste_Ccv_media | 42.38 |
| desgaste_Ccv_mediana | 43.54 |
| desgaste_Ccv_desviacion | 27.99 |
| desgaste_Ccv_min | 0.01 |
| desgaste_Ccv_max | 90.95 |
| desgaste_Ccv_IQR | 46.89 |
| tasa_desgaste_Ccv_acum_media | 3.62 |
| tasa_desgaste_Ccv_acum_mediana | 4.04 |
| tasa_desgaste_Ccv_acum_desviacion | 1.42 |
| tasa_desgaste_Ccv_acum_min | 0.01 |
| tasa_desgaste_Ccv_acum_max | 4.85 |
| tasa_desgaste_Ccv_acum_IQR | 0.47 |
| Shift_aplan_Ccv_media | 2.44 |
| Shift_aplan_Ccv_mediana | 2.56 |
| Shift_aplan_Ccv_desviacion | 0.65 |
| Shift_aplan_Ccv_min | 1.00 |
| Shift_aplan_Ccv_max | 3.40 |
| Shift_aplan_Ccv_IQR | 0.90 |
str(datosOpe5)
## 'data.frame': 32 obs. of 17 variables:
## $ Fecha : Date, format: "2019-01-20" "2019-02-13" ...
## $ log_T_H : num 11.4 11.5 11.4 11.4 11.4 ...
## $ log_Posion_poste_manto : num 87.3 114.8 170.1 107 177.8 ...
## $ log_CSS : num 2.01 2.01 2.01 2.01 2.01 ...
## $ log_Presion_Hidroset : num 263 275 252 247 245 ...
## $ log_Corriente : num 104.1 110.2 101.5 95.5 98.5 ...
## $ log_T_aceite_motor : num 3.89 3.93 3.95 3.95 3.94 ...
## $ log_F80_CV_001 : num 1.014 1.047 1.005 0.928 1.052 ...
## $ log_F80_CV_701 : num 0.891 0.926 1.091 1.08 0.954 ...
## $ log_F80_CV_005 : num 2.94 2.94 3.13 3.11 2.81 ...
## $ log_ton_acum : num 12834383 14997962 17028676 1308866 2804110 ...
## $ log_ton_acum2 : num 1.28e+07 1.50e+07 1.80e+07 1.00e-01 2.80e+06 ...
## $ log_ratio_presion_corriente: num 0.925 0.915 0.909 0.949 0.911 ...
## $ log_ratio_carga_corriente : num 6.79 6.77 6.78 6.85 6.82 ...
## $ log_ratio_mto_css : num 11.6 15.3 22.7 14.3 23.7 ...
## $ log_ratio_th_mto : num 1061 839 524 845 507 ...
## $ log_ratio_th_cp : num 5.53 5.52 5.53 5.58 5.57 ...
str(datosRepo5)
## tibble [32 × 5] (S3: tbl_df/tbl/data.frame)
## $ date_Time : Date[1:32], format: "2019-01-20" "2019-02-13" ...
## $ diametro_Mto : chr [1:32] "115" "115" "115" "110" ...
## $ desgaste_Ccv : num [1:32] 58.15 65.3 76.33 0.01 13.65 ...
## $ tasa_desgaste_Ccv_acum: num [1:32] 4.55 4.33 4.25 0.01 4.85 ...
## $ Shift_aplan_Ccv : num [1:32] 2.4 2.38 2.25 2.4 2.9 ...
# Definir una función para contar valores nulos o NaN en cada columna de un dataframe
count_missing <- function(df, df_name) {
df %>%
summarise_all(~ sum(is.na(.))) %>%
pivot_longer(cols = everything(), names_to = "variable", values_to = "count") %>%
mutate(issue_type = "missing",
dataset = df_name)
}
# Aplicar la función a los dataframes datosRepo5 y datosOpe5
repo_missing_counts <- count_missing(datosRepo5, "datosRepo5")
ope_missing_counts <- count_missing(datosOpe5, "datosOpe5")
missing_counts <- bind_rows(repo_missing_counts, ope_missing_counts)
kable(missing_counts, col.names = c("Variable", "Count", "Issue Type", "Dataset"), caption = "Counts of Missing Observations")
| Variable | Count | Issue Type | Dataset |
|---|---|---|---|
| date_Time | 0 | missing | datosRepo5 |
| diametro_Mto | 0 | missing | datosRepo5 |
| desgaste_Ccv | 0 | missing | datosRepo5 |
| tasa_desgaste_Ccv_acum | 0 | missing | datosRepo5 |
| Shift_aplan_Ccv | 0 | missing | datosRepo5 |
| Fecha | 0 | missing | datosOpe5 |
| log_T_H | 0 | missing | datosOpe5 |
| log_Posion_poste_manto | 0 | missing | datosOpe5 |
| log_CSS | 0 | missing | datosOpe5 |
| log_Presion_Hidroset | 0 | missing | datosOpe5 |
| log_Corriente | 0 | missing | datosOpe5 |
| log_T_aceite_motor | 0 | missing | datosOpe5 |
| log_F80_CV_001 | 0 | missing | datosOpe5 |
| log_F80_CV_701 | 0 | missing | datosOpe5 |
| log_F80_CV_005 | 0 | missing | datosOpe5 |
| log_ton_acum | 0 | missing | datosOpe5 |
| log_ton_acum2 | 0 | missing | datosOpe5 |
| log_ratio_presion_corriente | 0 | missing | datosOpe5 |
| log_ratio_carga_corriente | 0 | missing | datosOpe5 |
| log_ratio_mto_css | 0 | missing | datosOpe5 |
| log_ratio_th_mto | 0 | missing | datosOpe5 |
| log_ratio_th_cp | 0 | missing | datosOpe5 |
# Definir una función para contar observaciones iguales o menores a 0 en cada columna de un dataframe
count_zero_or_less <- function(df, df_name) {
df %>%
summarise_all(~ sum(. <= 0, na.rm = TRUE)) %>%
pivot_longer(cols = everything(), names_to = "variable", values_to = "count") %>%
mutate(issue_type = "menor o igual a 0",
dataset = df_name)
}
# Aplicar la función a los dataframes datosRepo5 y datosOpe5
repo_zero_counts <- count_zero_or_less(datosRepo5, "datosRepo5")
ope_zero_counts <- count_zero_or_less(datosOpe5, "datosOpe5")
# Unir los resultados
zero_counts <- bind_rows(repo_zero_counts, ope_zero_counts)
kable(zero_counts, col.names = c("Variable", "Count", "Issue Type", "Dataset"), caption = "Counts of Observations Equal to or Less Than 0")
| Variable | Count | Issue Type | Dataset |
|---|---|---|---|
| date_Time | 0 | menor o igual a 0 | datosRepo5 |
| diametro_Mto | 0 | menor o igual a 0 | datosRepo5 |
| desgaste_Ccv | 0 | menor o igual a 0 | datosRepo5 |
| tasa_desgaste_Ccv_acum | 0 | menor o igual a 0 | datosRepo5 |
| Shift_aplan_Ccv | 0 | menor o igual a 0 | datosRepo5 |
| Fecha | 0 | menor o igual a 0 | datosOpe5 |
| log_T_H | 0 | menor o igual a 0 | datosOpe5 |
| log_Posion_poste_manto | 0 | menor o igual a 0 | datosOpe5 |
| log_CSS | 0 | menor o igual a 0 | datosOpe5 |
| log_Presion_Hidroset | 0 | menor o igual a 0 | datosOpe5 |
| log_Corriente | 0 | menor o igual a 0 | datosOpe5 |
| log_T_aceite_motor | 0 | menor o igual a 0 | datosOpe5 |
| log_F80_CV_001 | 0 | menor o igual a 0 | datosOpe5 |
| log_F80_CV_701 | 0 | menor o igual a 0 | datosOpe5 |
| log_F80_CV_005 | 0 | menor o igual a 0 | datosOpe5 |
| log_ton_acum | 0 | menor o igual a 0 | datosOpe5 |
| log_ton_acum2 | 0 | menor o igual a 0 | datosOpe5 |
| log_ratio_presion_corriente | 0 | menor o igual a 0 | datosOpe5 |
| log_ratio_carga_corriente | 0 | menor o igual a 0 | datosOpe5 |
| log_ratio_mto_css | 0 | menor o igual a 0 | datosOpe5 |
| log_ratio_th_mto | 0 | menor o igual a 0 | datosOpe5 |
| log_ratio_th_cp | 0 | menor o igual a 0 | datosOpe5 |
# Filtrar las observaciones donde log_ton_acum2 es igual o menor a 0
observaciones_iguales_o_menores_a_cero <- datosOpe5 %>%
filter(log_ton_acum2 <= 0)
# Mostrar las observaciones capturadas
observaciones_iguales_o_menores_a_cero
## [1] Fecha log_T_H
## [3] log_Posion_poste_manto log_CSS
## [5] log_Presion_Hidroset log_Corriente
## [7] log_T_aceite_motor log_F80_CV_001
## [9] log_F80_CV_701 log_F80_CV_005
## [11] log_ton_acum log_ton_acum2
## [13] log_ratio_presion_corriente log_ratio_carga_corriente
## [15] log_ratio_mto_css log_ratio_th_mto
## [17] log_ratio_th_cp
## <0 rows> (o 0- extensión row.names)
# Calcular la mediana de log_ton_acum2 excluyendo los valores negativos
mediana_log_ton_acum2 <- datosOpe5 %>%
filter(log_ton_acum2 > 0) %>%
summarise(mediana = median(log_ton_acum2, na.rm = TRUE)) %>%
pull(mediana)
# Reemplazar los valores negativos por la mediana
datosOpe5 <- datosOpe5 %>%
mutate(log_ton_acum2 = ifelse(log_ton_acum2 <= 0, mediana_log_ton_acum2, log_ton_acum2))
# Crea el dataframe con los datos proporcionados para la primera tabla
tabla_ccv <- data.frame(
"Fila de ccv" = c("Superior", "Intermedia", "Intermedia", "Inferior", "", ""),
"Volumen por pieza m3" = c(0.0744375, 0.065125, 0.05675, 0.0708125, "", ""),
"Volumen por fila m3" = c(1.191, 1.042, 0.908, 1.133, "", ""),
"Peso por pieza Ton" = c(0.584334375, 0.51123125, 0.4454875, 0.555878125, "Total", ""),
"Peso por fila Ton" = c(9.34935, 8.1797, 7.1278, 8.89405, 33.5509, "Para el cálculo del peso se considera una densidad del acero de 7,85 Ton/m3")
)
# Convierte la primera tabla a formato kable y aplica estilos
tabla_kable_ccv <- kable(tabla_ccv, format = "html", align = c("c", "c", "c", "c", "c"), escape = F) %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive"), full_width = F, position = "center") %>%
row_spec(5, bold = T, color = "white", background = "gray") %>%
row_spec(6, italic = T)
# Imprime la primera tabla
print(tabla_kable_ccv)
## <table class="table table-striped table-hover table-condensed table-responsive" style="color: black; width: auto !important; margin-left: auto; margin-right: auto;">
## <thead>
## <tr>
## <th style="text-align:center;"> Fila.de.ccv </th>
## <th style="text-align:center;"> Volumen.por.pieza.m3 </th>
## <th style="text-align:center;"> Volumen.por.fila.m3 </th>
## <th style="text-align:center;"> Peso.por.pieza.Ton </th>
## <th style="text-align:center;"> Peso.por.fila.Ton </th>
## </tr>
## </thead>
## <tbody>
## <tr>
## <td style="text-align:center;"> Superior </td>
## <td style="text-align:center;"> 0.0744375 </td>
## <td style="text-align:center;"> 1.191 </td>
## <td style="text-align:center;"> 0.584334375 </td>
## <td style="text-align:center;"> 9.34935 </td>
## </tr>
## <tr>
## <td style="text-align:center;"> Intermedia </td>
## <td style="text-align:center;"> 0.065125 </td>
## <td style="text-align:center;"> 1.042 </td>
## <td style="text-align:center;"> 0.51123125 </td>
## <td style="text-align:center;"> 8.1797 </td>
## </tr>
## <tr>
## <td style="text-align:center;"> Intermedia </td>
## <td style="text-align:center;"> 0.05675 </td>
## <td style="text-align:center;"> 0.908 </td>
## <td style="text-align:center;"> 0.4454875 </td>
## <td style="text-align:center;"> 7.1278 </td>
## </tr>
## <tr>
## <td style="text-align:center;"> Inferior </td>
## <td style="text-align:center;"> 0.0708125 </td>
## <td style="text-align:center;"> 1.133 </td>
## <td style="text-align:center;"> 0.555878125 </td>
## <td style="text-align:center;"> 8.89405 </td>
## </tr>
## <tr>
## <td style="text-align:center;font-weight: bold;color: white !important;background-color: gray !important;"> </td>
## <td style="text-align:center;font-weight: bold;color: white !important;background-color: gray !important;"> </td>
## <td style="text-align:center;font-weight: bold;color: white !important;background-color: gray !important;"> </td>
## <td style="text-align:center;font-weight: bold;color: white !important;background-color: gray !important;"> Total </td>
## <td style="text-align:center;font-weight: bold;color: white !important;background-color: gray !important;"> 33.5509 </td>
## </tr>
## <tr>
## <td style="text-align:center;font-style: italic;"> </td>
## <td style="text-align:center;font-style: italic;"> </td>
## <td style="text-align:center;font-style: italic;"> </td>
## <td style="text-align:center;font-style: italic;"> </td>
## <td style="text-align:center;font-style: italic;"> Para el cálculo del peso se considera una densidad del acero de 7,85 Ton/m3 </td>
## </tr>
## </tbody>
## </table>
# Crea el dataframe con los datos proporcionados para la segunda tabla
tabla_acero <- data.frame(
"Kilogramos de acero en cóncavas" = 33550.90,
"Cantidad de piezas (16 x 4 filas)" = 64,
"Peso por pieza individual aprox." = 524.23,
"Precio de USD/Kg aprox." = 18,
"Precio por pieza (USD) aprox." = 9436.19
)
# Convierte la segunda tabla a formato kable y aplica estilos
tabla_kable_acero <- kable(tabla_acero, format = "html", align = c("c", "c", "c", "c", "c"), escape = F) %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive"), full_width = F, position = "center")
# Imprime la segunda tabla
print(tabla_kable_acero)
## <table class="table table-striped table-hover table-condensed table-responsive" style="color: black; width: auto !important; margin-left: auto; margin-right: auto;">
## <thead>
## <tr>
## <th style="text-align:center;"> Kilogramos.de.acero.en.cóncavas </th>
## <th style="text-align:center;"> Cantidad.de.piezas..16.x.4.filas. </th>
## <th style="text-align:center;"> Peso.por.pieza.individual.aprox. </th>
## <th style="text-align:center;"> Precio.de.USD.Kg.aprox. </th>
## <th style="text-align:center;"> Precio.por.pieza..USD..aprox. </th>
## </tr>
## </thead>
## <tbody>
## <tr>
## <td style="text-align:center;"> 33550.9 </td>
## <td style="text-align:center;"> 64 </td>
## <td style="text-align:center;"> 524.23 </td>
## <td style="text-align:center;"> 18 </td>
## <td style="text-align:center;"> 9436.19 </td>
## </tr>
## </tbody>
## </table>
# Crea el dataframe con los datos adicionales para la tercera tabla
tabla_adicional <- data.frame(
"Concepto" = c("Precio de cóncavas", "Costo de parada no programada por hora", "Umbral de desgaste de cóncavas"),
"Valor" = c("10000 USD", "100000 USD", "80 mm")
)
# Convierte la tercera tabla a formato kable y aplica estilos
tabla_kable_adicional <- kable(tabla_adicional, format = "html", align = c("c", "c"), escape = F) %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive"), full_width = F, position = "center")
# Imprime la tercera tabla
print(tabla_kable_adicional)
## <table class="table table-striped table-hover table-condensed table-responsive" style="color: black; width: auto !important; margin-left: auto; margin-right: auto;">
## <thead>
## <tr>
## <th style="text-align:center;"> Concepto </th>
## <th style="text-align:center;"> Valor </th>
## </tr>
## </thead>
## <tbody>
## <tr>
## <td style="text-align:center;"> Precio de cóncavas </td>
## <td style="text-align:center;"> 10000 USD </td>
## </tr>
## <tr>
## <td style="text-align:center;"> Costo de parada no programada por hora </td>
## <td style="text-align:center;"> 100000 USD </td>
## </tr>
## <tr>
## <td style="text-align:center;"> Umbral de desgaste de cóncavas </td>
## <td style="text-align:center;"> 80 mm </td>
## </tr>
## </tbody>
## </table>
datosOpe6 <- datosOpe5
datosRepo6 <- datosRepo5
# Unir los datasets por la columna 'Fecha'
datosCombinados <- merge(datosOpe6, datosRepo6, by.x = "Fecha", by.y = "date_Time")
# Contar valores NaN o nulos
missing_counts2 <- count_missing(datosCombinados, "datosCombinados")
kable(missing_counts2, col.names = c("Variable", "Count", "Issue Type", "Dataset"), caption = "Counts of Missing Observations")
| Variable | Count | Issue Type | Dataset |
|---|---|---|---|
| Fecha | 0 | missing | datosCombinados |
| log_T_H | 0 | missing | datosCombinados |
| log_Posion_poste_manto | 0 | missing | datosCombinados |
| log_CSS | 0 | missing | datosCombinados |
| log_Presion_Hidroset | 0 | missing | datosCombinados |
| log_Corriente | 0 | missing | datosCombinados |
| log_T_aceite_motor | 0 | missing | datosCombinados |
| log_F80_CV_001 | 0 | missing | datosCombinados |
| log_F80_CV_701 | 0 | missing | datosCombinados |
| log_F80_CV_005 | 0 | missing | datosCombinados |
| log_ton_acum | 0 | missing | datosCombinados |
| log_ton_acum2 | 0 | missing | datosCombinados |
| log_ratio_presion_corriente | 0 | missing | datosCombinados |
| log_ratio_carga_corriente | 0 | missing | datosCombinados |
| log_ratio_mto_css | 0 | missing | datosCombinados |
| log_ratio_th_mto | 0 | missing | datosCombinados |
| log_ratio_th_cp | 0 | missing | datosCombinados |
| diametro_Mto | 0 | missing | datosCombinados |
| desgaste_Ccv | 0 | missing | datosCombinados |
| tasa_desgaste_Ccv_acum | 0 | missing | datosCombinados |
| Shift_aplan_Ccv | 0 | missing | datosCombinados |
# Contar los valores no positivos ( X <= 0)
zero_counts2 <- count_zero_or_less(datosCombinados, "datosCombinados")
kable(zero_counts2, col.names = c("Variable", "Count", "Issue Type", "Dataset"), caption = "Counts of Observations Equal to or Less Than 0")
| Variable | Count | Issue Type | Dataset |
|---|---|---|---|
| Fecha | 0 | menor o igual a 0 | datosCombinados |
| log_T_H | 0 | menor o igual a 0 | datosCombinados |
| log_Posion_poste_manto | 0 | menor o igual a 0 | datosCombinados |
| log_CSS | 0 | menor o igual a 0 | datosCombinados |
| log_Presion_Hidroset | 0 | menor o igual a 0 | datosCombinados |
| log_Corriente | 0 | menor o igual a 0 | datosCombinados |
| log_T_aceite_motor | 0 | menor o igual a 0 | datosCombinados |
| log_F80_CV_001 | 0 | menor o igual a 0 | datosCombinados |
| log_F80_CV_701 | 0 | menor o igual a 0 | datosCombinados |
| log_F80_CV_005 | 0 | menor o igual a 0 | datosCombinados |
| log_ton_acum | 0 | menor o igual a 0 | datosCombinados |
| log_ton_acum2 | 0 | menor o igual a 0 | datosCombinados |
| log_ratio_presion_corriente | 0 | menor o igual a 0 | datosCombinados |
| log_ratio_carga_corriente | 0 | menor o igual a 0 | datosCombinados |
| log_ratio_mto_css | 0 | menor o igual a 0 | datosCombinados |
| log_ratio_th_mto | 0 | menor o igual a 0 | datosCombinados |
| log_ratio_th_cp | 0 | menor o igual a 0 | datosCombinados |
| diametro_Mto | 0 | menor o igual a 0 | datosCombinados |
| desgaste_Ccv | 0 | menor o igual a 0 | datosCombinados |
| tasa_desgaste_Ccv_acum | 0 | menor o igual a 0 | datosCombinados |
| Shift_aplan_Ccv | 0 | menor o igual a 0 | datosCombinados |
# Tratar la columna 'diametro_Mto' mediante One-Hot Encoding
datosCombinados <- datosCombinados %>%
mutate(diametro_Mto = as.factor(diametro_Mto)) %>%
tidyr::pivot_wider(names_from = diametro_Mto, values_from = diametro_Mto, values_fill = 0, values_fn = length) %>%
rename_with(~ paste0("diametro_", .), starts_with("115"), starts_with("110"))
set.seed(19831118) # Semilla para reproducibilidad
index <- createDataPartition(datosCombinados$desgaste_Ccv, p = 0.8, list = FALSE)
train_set <- datosCombinados[index,]
test_set <- datosCombinados[-index,]
# Ajustar el modelo lineal
modelo <- lm(desgaste_Ccv ~ ., data = train_set)
# Resumen del modelo
summary(modelo)
##
## Call:
## lm(formula = desgaste_Ccv ~ ., data = train_set)
##
## Residuals:
## 1 2 3 4 5 6 7
## 0.5146108 -2.1164401 -1.2606466 1.3635760 0.3401425 0.5264404 1.7697118
## 8 9 10 11 12 13 14
## 2.1190625 -1.1135561 -0.6794954 0.1935941 -1.0375869 1.2551572 -0.4111645
## 15 16 17 18 19 20 21
## -2.6710782 0.1416747 0.2216282 -0.4896750 -0.3588343 0.3954667 0.8722560
## 22 23 24 25 26 27 28
## -0.0253725 0.0008771 0.2609189 -0.8672741 0.5859146 0.3430425 0.1270496
##
## Coefficients: (1 not defined because of singularities)
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) -3.648e+03 8.280e+03 -0.441 0.678
## Fecha -3.416e-03 7.344e-03 -0.465 0.661
## log_T_H 5.752e+02 8.799e+02 0.654 0.542
## log_Posion_poste_manto -2.933e-01 1.317e+00 -0.223 0.833
## log_CSS 9.913e+01 3.178e+02 0.312 0.768
## log_Presion_Hidroset -2.132e+00 3.014e+00 -0.707 0.511
## log_Corriente -2.991e-01 5.757e+00 -0.052 0.961
## log_T_aceite_motor -3.955e+01 3.324e+01 -1.190 0.288
## log_F80_CV_001 1.097e+00 2.683e+01 0.041 0.969
## log_F80_CV_701 3.843e+00 1.447e+01 0.266 0.801
## log_F80_CV_005 9.919e-01 2.527e+00 0.393 0.711
## log_ton_acum 4.073e-06 5.532e-06 0.736 0.495
## log_ton_acum2 -8.766e-08 5.180e-06 -0.017 0.987
## log_ratio_presion_corriente -7.074e+02 5.825e+03 -0.121 0.908
## log_ratio_carga_corriente 1.214e+03 7.710e+03 0.157 0.881
## log_ratio_mto_css 1.773e+00 1.029e+01 0.172 0.870
## log_ratio_th_mto -1.161e-02 1.080e-02 -1.076 0.331
## log_ratio_th_cp -1.791e+03 8.223e+03 -0.218 0.836
## tasa_desgaste_Ccv_acum 1.802e+00 1.478e+00 1.219 0.277
## Shift_aplan_Ccv -3.231e-01 3.749e+00 -0.086 0.935
## diametro_115 4.710e+00 8.743e+00 0.539 0.613
## `110` 3.238e+00 6.397e+00 0.506 0.634
## `113` 5.468e+00 1.022e+01 0.535 0.616
## `113_2` NA NA NA NA
##
## Residual standard error: 2.472 on 5 degrees of freedom
## Multiple R-squared: 0.9986, Adjusted R-squared: 0.9922
## F-statistic: 157.8 on 22 and 5 DF, p-value: 1.101e-05
# Predecir en el conjunto de prueba
predicciones <- predict(modelo, newdata = test_set)
# Calcular MSE y R²
mse_test <- mse(test_set$desgaste_Ccv, predicciones)
r2_test <- R2(test_set$desgaste_Ccv, predicciones)
# Mostrar los resultados de MSE y R²
print(paste("MSE en el conjunto de prueba:", mse_test))
## [1] "MSE en el conjunto de prueba: 6.05658857717856"
print(paste("R² en el conjunto de prueba:", r2_test))
## [1] "R² en el conjunto de prueba: 0.996187890407707"
# Opcional: Gráfico de valores observados vs. predichos
plot(test_set$desgaste_Ccv, predicciones, main = "Valores Observados vs. Predichos",
xlab = "Valores Observados", ylab = "Valores Predichos")
abline(0, 1, col = "red")
# Función para simular el desgaste bajo diferentes escenarios operacionales
simulate_wear <- function(data, adjustments) {
simulated_data <- data
for (var in names(adjustments)) {
simulated_data[[var]] <- simulated_data[[var]] * adjustments[[var]]
}
predicted_wear <- predict(modelo, newdata = simulated_data)
return(predicted_wear)
}
# Definir ajustes en las variables operacionales (por ejemplo, reducir/incrementar en 10%)
adjustments <- list("log_Presion_Hidroset" = 0.9, "log_Corriente" = 1.1)
# Simular el desgaste bajo estos ajustes
predicted_wear_adjusted <- simulate_wear(test_set, adjustments)
# Comparar desgaste predicho original vs. ajustado
mean_original <- mean(test_set$desgaste_Ccv)
mean_adjusted <- mean(predicted_wear_adjusted)
percentage_change <- ((mean_adjusted - mean_original) / mean_original) * 100
# Análisis de costo-beneficio
additional_weeks <- (mean_adjusted - 80) / mean(test_set$tasa_desgaste_Ccv_acum)
cost_savings <- additional_weeks * 100000 - 10000
# Gráfico de desgaste original vs. ajustado
plot(test_set$Fecha, test_set$desgaste_Ccv, type = "o", col = "blue", ylim = c(0, max(c(test_set$desgaste_Ccv, predicted_wear_adjusted))), xlab = "Fecha", ylab = "Desgaste de Ccv", main = "Comparación de Desgaste Ccv: Original vs. Ajustado")
lines(test_set$Fecha, predicted_wear_adjusted, type = "o", col = "red")
legend("topright", legend = c("Original", "Ajustado"), col = c("blue", "red"), lty = 1)
# Crear un data frame con los resultados redondeados
resultados <- data.frame(
Descripcion = c( "Semanas adicionales antes de alcanzar el umbral de 80 mm", "Ahorro estimado en costos (USD)"),
Valor = c(sprintf("%.2f", additional_weeks), sprintf("$%.2f", cost_savings))
)
# Generar tabla kable y convertirla en un objeto gráfico
resultados_kable <- kable(resultados, "html", col.names = c("Descripción", "Valor"), align = c('l', 'c')) %>%
kable_styling(bootstrap_options = c("striped", "hover", "condensed", "responsive"), full_width = F, position = "center") %>%
column_spec(1, bold = TRUE, width = "30em") %>%
column_spec(2, width = "20em")
# Imprimir la tabla kable
print(resultados_kable)
## <table class="table table-striped table-hover table-condensed table-responsive" style="color: black; width: auto !important; margin-left: auto; margin-right: auto;">
## <thead>
## <tr>
## <th style="text-align:left;"> Descripción </th>
## <th style="text-align:center;"> Valor </th>
## </tr>
## </thead>
## <tbody>
## <tr>
## <td style="text-align:left;width: 30em; font-weight: bold;"> Semanas adicionales antes de alcanzar el umbral de 80 mm </td>
## <td style="text-align:center;width: 20em; "> 2.86 </td>
## </tr>
## <tr>
## <td style="text-align:left;width: 30em; font-weight: bold;"> Ahorro estimado en costos (USD) </td>
## <td style="text-align:center;width: 20em; "> $275730.14 </td>
## </tr>
## </tbody>
## </table>