From 992d60336a35bc1935b16377f77eaaa329d6b06d Mon Sep 17 00:00:00 2001 From: Johannes Findeisen Date: Wed, 19 Oct 2022 00:54:19 +0200 Subject: [PATCH] Initial commit of the first prototype. --- .gitignore | 5 +- LICENSE | 2 +- Makefile | 2 + README.md | 16 ++- idea.mmd | 17 --- man2book | 74 ++++++++++++ output/.gitkeep | 0 requirements.txt | 2 + tmp/df.html | 298 ----------------------------------------------- workflow.txt | 0 10 files changed, 97 insertions(+), 319 deletions(-) create mode 100644 Makefile delete mode 100644 idea.mmd create mode 100755 man2book create mode 100644 output/.gitkeep create mode 100644 requirements.txt delete mode 100644 tmp/df.html create mode 100644 workflow.txt diff --git a/.gitignore b/.gitignore index 309adef..35c20c8 100644 --- a/.gitignore +++ b/.gitignore @@ -1,2 +1,5 @@ .idea/ -man2ebook.iml +man2book.iml +/output/*.html +/tmp/ +/venv/ \ No newline at end of file diff --git a/LICENSE b/LICENSE index d449d3e..7f26d00 100644 --- a/LICENSE +++ b/LICENSE @@ -1,6 +1,6 @@ MIT License -Copyright (c) +Copyright (c) 2022 Johannes Findeisen Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal diff --git a/Makefile b/Makefile new file mode 100644 index 0000000..b91b5cd --- /dev/null +++ b/Makefile @@ -0,0 +1,2 @@ +clean: + rm -rf ./output/*.html diff --git a/README.md b/README.md index 1f55875..f4c7103 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ -# man2ebook +# man2book -A tool to convert all installed man pages to a simple ebook with contents in the head. +A tool to convert all installed man pages to a simple ebook or other formats with contents in the head. Maybe this could be only a small shell script but if not I will use Python. @@ -24,3 +24,15 @@ Not more... ;) - https://docutils.sourceforge.io/ - Python (as I can see now it only converts from .rst files.) - https://github.com/hanez/aov-html2epub - Bash + +## Workflow + +1. Read manpage sections +2. Create HTML file for each manpage in each section +3. Read and store title(metadata) for in each file +4. Remove all unneeded stuff like html, head and body in each file +5. Create TOC for the entire ebook +6. Create a header HTML file for the ebook containing the head tag +7. Merge toc and content of all files for preparing the ebook +8. Merge the header file with the prepared ebook while inserting html and body tags which should result in a valid HTML file +9. Create ebook from the resulting HTML file using Pandoc diff --git a/idea.mmd b/idea.mmd deleted file mode 100644 index b1c49c4..0000000 --- a/idea.mmd +++ /dev/null @@ -1,17 +0,0 @@ -Mind Map generated by NB MindMap plugin -> __version__=`1.1`,showJumps=`true` ---- - -# man page - -## "pandoc", "man2html" or "man \-Thtml df \> df\.html" - -### maybe create a metadata file - -### create contents and index files - -### cleanup contents and index files - -### merge contents and index files - -### create epub file \(pandoc\) diff --git a/man2book b/man2book new file mode 100755 index 0000000..1a8d05b --- /dev/null +++ b/man2book @@ -0,0 +1,74 @@ +#!/usr/bin/python3 +import argparse +import os + +# the limit is just for development to limit the number of man pages in each section. set to 0 to +# have no limit. +LIMIT = 0 +MANPAGE_PATH = '/usr/share/man/' +OUTPUT_DIR = '~/code/man2book/output/' +TMP_DIR = '/tmp' + + +__author__ = 'Johannes Findeisen ' + OUTPUT_DIR + 'man' + section + + '.' + manpage + '.html') + if LIMIT > 0: + if x == LIMIT: + break + x += 1 diff --git a/output/.gitkeep b/output/.gitkeep new file mode 100644 index 0000000..e69de29 diff --git a/requirements.txt b/requirements.txt new file mode 100644 index 0000000..5632f4c --- /dev/null +++ b/requirements.txt @@ -0,0 +1,2 @@ +lxml~=4.9.1 +pandocfilters~=1.5.0 \ No newline at end of file diff --git a/tmp/df.html b/tmp/df.html deleted file mode 100644 index a6150a5..0000000 --- a/tmp/df.html +++ /dev/null @@ -1,298 +0,0 @@ - - - - - - - - - -DF - - - - -

DF

- -NAME
-SYNOPSIS
-DESCRIPTION
-OPTIONS
-AUTHOR
-REPORTING BUGS
-COPYRIGHT
-SEE ALSO
- -
- - -

NAME - -

- - -

df − -report file system space usage

- -

SYNOPSIS - -

- - -

df -[OPTION]... [FILE]...

- -

DESCRIPTION - -

- - -

This manual -page documents the GNU version of df. df -displays the amount of space available on the file system -containing each file name argument. If no file name is -given, the space available on all currently mounted file -systems is shown. Space is shown in 1K blocks by default, -unless the environment variable POSIXLY_CORRECT is set, in -which case 512-byte blocks are used.

- -

If an argument -is the absolute file name of a device node containing a -mounted file system, df shows the space available on -that file system rather than on the file system containing -the device node. This version of df cannot show the -space available on unmounted file systems, because on most -kinds of systems doing so requires very nonportable intimate -knowledge of file system structures.

- -

OPTIONS - -

- - -

Show -information about the file system on which each FILE -resides, or all file systems by default.

- -

Mandatory -arguments to long options are mandatory for short options -too.
-−a
, −−all

- -

include pseudo, duplicate, -inaccessible file systems

- -

−B, -−−block−size=SIZE

- -

scale sizes by SIZE before -printing them; e.g., ’−BM’ prints sizes in -units of 1,048,576 bytes; see SIZE format below

- -

−h, -−−human−readable

- -

print sizes in powers of 1024 -(e.g., 1023M)

- -

−H, -−−si

- -

print sizes in powers of 1000 -(e.g., 1.1G)

- -

−i, -−−inodes

- -

list inode information instead -of block usage

- - - - - - - - -
- - -

−k

- - -

like −−block−size=1K

-
- -

−l, -−−local

- -

limit listing to local file -systems

- - -

−−no−sync

- -

do not invoke sync before -getting usage info (default)

- - -

−−output[=FIELD_LIST]

- -

use the output format defined -by FIELD_LIST, or print all fields if FIELD_LIST is -omitted.

- -

−P, -−−portability

- -

use the POSIX output format

- - - - - - - - -
- - -

−−sync

- - -

invoke sync before getting usage info

-
- -

−−total

- -

elide all entries insignificant -to available space, and produce a grand total

- -

−t, -−−type=TYPE

- -

limit listing to file systems -of type TYPE

- -

−T, -−−print−type

- -

print file system type

- -

−x, -−−exclude−type=TYPE

- -

limit listing to file systems -not of type TYPE

- - - - - - - - - - - - - - -
- - -

−v

- - -

(ignored)

-
- - -

−−help

- - -

display this help and exit

-
- - -

−−version

- -

output version information and -exit

- -

Display values -are in units of the first available SIZE from -−−block−size, and the -DF_BLOCK_SIZE, BLOCK_SIZE and BLOCKSIZE environment -variables. Otherwise, units default to 1024 bytes (or 512 if -POSIXLY_CORRECT is set).

- -

The SIZE -argument is an integer and optional unit (example: 10K is -10*1024). Units are K,M,G,T,P,E,Z,Y (powers of 1024) or -KB,MB,... (powers of 1000). Binary prefixes can be used, -too: KiB=K, MiB=M, and so on.

- -

FIELD_LIST is a -comma−separated list of columns to be included. Valid -field names are: ’source’, ’fstype’, -’itotal’, ’iused’, -’iavail’, ’ipcent’, -’size’, ’used’, ’avail’, -’pcent’, ’file’ and -’target’ (see info page).

- -

AUTHOR - -

- - -

Written by -Torbjorn Granlund, David MacKenzie, and Paul Eggert.

- -

REPORTING BUGS - -

- - -

GNU coreutils -online help: <https://www.gnu.org/software/coreutils/> -
-Report any translation bugs to -<https://translationproject.org/team/>

- -

COPYRIGHT - -

- - -

Copyright -© 2022 Free Software Foundation, Inc. License GPLv3+: -GNU GPL version 3 or later -<https://gnu.org/licenses/gpl.html>.
-This is free software: you are free to change and -redistribute it. There is NO WARRANTY, to the extent -permitted by law.

- -

SEE ALSO - -

- - -

Full -documentation -<https://www.gnu.org/software/coreutils/df>
-or available locally via: info '(coreutils) df -invocation'

-
- - diff --git a/workflow.txt b/workflow.txt new file mode 100644 index 0000000..e69de29